Rotary Position Embedding for Vision Transformer
Byeongho Heo, Song Park, Dongyoon Han, Sangdoo Yun
Introduction
Transformers have become popular in neural architecture due to their strong performance across various tasks in language and computer vision domains . The transformer treats input data as a sequence of tokens. The tokens equally interact with others through a self-attention mechanism . Since the self-attention mechanism is independent of the token index or positions (i.e., permutation invariance), the transformer requires additional position information, usually injected by position embedding . The position embeddings give the position information to input tokens with specific embedding designed for the transformer. They uniquely differentiate tokens based on their locations rather than their contents. Thus, the position information of self-attention heavily depends on the position embedding, which is a crucial component in designing transformer architectures.
There are two primary methods in position embedding for Vision Transformers: Absolute Positional Embedding (APE) and Relative Position Bias (RPB) . APE utilizes the absolute position of tokens for position embedding through sinusoidal or learnable embedding. Otherwise, RPB enables relative positions between tokens by adding relative biases to the attention matrix of the self-attention layers. In general, APE is used for traditional ViT architecture , and RPE is preferred to hierarchical ViT like Swin Transformer . Although both position embeddings are effective for the transformer on fixed-resolution settings, they struggle with resolution changes, requiring flexibility and extrapolation in position embeddings. Considering that the resolution of pre-training is usually smaller than that of downstream dense prediction, it might degrade ViT performance in various applications, such as multi-resolution recognition, object detection, and segmentation.
This paper aims to improve position embedding for vision transformers by applying an extended Rotary Position Embedding (RoPE) . RoPE is a relative position embedding that is specially designed for extrapolation in language domains. Despite the remarkable success of RoPE in Large-Language Models , its effectiveness in vision tasks has not been validated due to limited investigation. In this paper, we provide a comprehensive investigation of RoPE for transformers in vision recognition tasks. Our investigation starts with 1D to 2D expansion of RoPE to cope with images rather than original language inputs. Although 2D RoPE using axial frequencies was used in pioneer works , we argue that it lacks the ability to handle diagonal directions, which are preferred in convolution networks by the square kernel. To cope with the diagonal direction of RoPE, we propose to use mixed axis frequencies for 2D RoPE, named RoPE-Mixed. Since RoPE-Mixed uses frequencies for both axes as learnable parameters of the network, it effectively handles diagonal direction and is more suitable for ViT’s attention than Axial 2D RoPE.
In experiments, we apply variants of 2D RoPE to representative transformer architectures, ViT and Swin Transformer, and validate the effects of 2D RoPE in various tasks, including multi-resolution classification on ImageNet-1k , object detection on MS-COCO , and semantic segmentation on ADE20k . The results show that 2D RoPE is a beneficial option as position embedding for transformers with an impressive performance improvement on high-resolution images, i.e., extrapolation of images. We believe our study demonstrates the impact of 2D RoPE in vision domains and contributes to future research by suggesting a beneficial option in position embedding for vision transformers.
Related Works
ViT introduces a transformer architecture for visual inputs, employing Absolute Positional Embedding (APE) . APE with learnable parameters effectively injects spatial positions of each token to be used for the self-attention mechanism. Hierarchical ViT such as Swin Transformer increase the spatial length of tokens at early layers using pooling. To handle a large number of tokens with limited position embeddings, Relative Position Bias (RPB) is preferred by the hierarchical ViTs. Studies have been conducted to improve position embedding for ViT based on these two major position embeddings. iRPE proposes an improved RPB by applying relative position embedding as multiplication with query vector. CPE finds that a convolution network can effectively inject relative position information to tokens and utilizes depth-wise convolution as conditional position embedding. LaPE shows that simple scaling with adaptive layer-norm can improve the positional embedding of various networks.
0.2 RoPE in vision modeling.
Pioneering studies introduced RoPE to ViT-related architectures. Hybrid X-former applies 1D RoPE to ViT variants named Vision X-formers; it is the first attempt at the application of RoPE in ViT to our knowledge. However, 1D RoPE is insufficient to demonstrate performance, and evaluation is limited to small datasets such as CIFAR and Tiny ImageNet . EVA-02 introduces 2D Axial RoPE to a new language-aligned vision model EVA-02, like CLIP . Unified-IO 2 uses 2D RoPE for new multi-modal modeling; 2D Axial RoPE is applied to non-text modalities, including vision, audio, and video. In diffusion modeling , FiT applies 2D Axial RoPE for their new diffusion model. In these studies, 2D Axial RoPE was used to improve new model performance on language-related or generation tasks, which differs from our goal of challenging classification, detection, and segmentation tasks. Exploring the impacts of 2D RoPE implementations in basic architectures with general training recipes could benefit diverse vision researchers.
0.3 Multi-resolution inference.
Diverged from Convolutional Neural Networks (ConvNets) , ViT requires a transformation in position embedding for multi-resolution inference. Some studies investigated a multi-resolution inference method for ViT to improve performance in downstream tasks. CAPE analyzes ViT’s position embedding in resolution changes and finds that augmenting position embedding improves the multi-resolution performance of ViT. Thus, they propose a training recipe with continuous augmenting of position embedding (CAPE). ResFormer shows that relative position embedding based on depth-wise convolution layer is beneficial for multi-resolution inference. Using this property, the study proposes an improved ViT architecture with global and local depth-wise conv embedding. It substantially improves multi-resolution performance with multi-resolution self-distillation learning recipes. In contrast to conventional multi-resolution, FlexiViT proposes a ViT with flexible patch sizes that can replace multi-resolution inference. In FlexiViT, ViT increases the patch size instead of increasing input resolution. By training with a multi-patch-size training scheme and distillation using ViT-B/8 , FlexiViT exhibits remarkable performance for various patch-size, which corresponds to multi-resolution in computation cost aspect.
These studies require special training methods, which make them difficult to combine with other training recipes, potentially reducing general applicability. RoPE improves multi-resolution performance while using existing training recipes as is, offering generally applicable and easy-to-use compared to others.
Method
Rotary Position Embedding (RoPE) was introduced to apply to key and query in self-attention layers as channel-wise multiplications, which is distinct from conventional position embeddings - APE is added to the stem layer; RPB is added to an attention matrix. We first present conventional position embeddings, including RoPE in language model at §3.1, and provide feasible expansion of RoPE to 2D inputs for transformers in the vision domain in subsequent §3.2. In §3.3, we describe the characteristics of RoPE compared to other position embedding and analysis for 2D RoPE.
Note that we use 0-base numbers for indexes , and . The other implementation of APE is to use learnable parameters and train them with the training process. learnable parameters are randomly initialized and are used as Eq. 1. It is the simplest way for APE, and supervised learning recipes use APE with learnable parameters . Since learnable APE is commonly used for ViT, we refer to it as the default option for APE.
1.2 Relative Position Bias (RPB)
Then, RPB is added to the attention matrix in Eq. 4 as
By RPB, self-attention handles relative positions. Note that we describe RPB for a head in a multi-head self-attention layer. Thus, in practice, RPB parameters and addition are repeated for each head in multi-head attention.
1.3 Rotary Position Embedding (RoPE)
is a recent method in the line of relative position embedding studies. Although RPB delivers relative position to the attention, simple addition as bias may limit interaction with attention weights, which causes limited utilization of relative position. Thus, RoFormer proposes a novel relative position embedding method: Rotary Position Embedding (RoPE). Note that this section explains the original RoPE designed for language modeling. We will explain our RoPE for 2D images in §3.2
Then, -th component of attention matrix is calculated as
and applied to query and key with the Hadamard product as
Note that the attention matrix with RoPE implies relative position in rotation form for () number of frequencies, which gives a lot of performance beneficial to the transformer, especially for extrapolation on inference stage based on periodic functions.
2 RoPE for 2D images
RoPE exhibits remarkable performance in the language domain. However, only a few studies have been conducted on using RoPE in the vision domain with 2D input, as it was designed solely for 1D input. This section introduces feasible implementations of 2D RoPE for image inputs: axial and learnable frequency.
A typical way to expand 1D position embedding to 2D is repeating 1D operation for each axis. Similar to 2D sinusoidal embedding in Eq. 2, axial frequency is to divide embedding dimensions into two and apply position embedding for the x-axis and y-axis separately. It is straightforward because it is technically the same as repeating 1D embedding twice.
Also, the range of position indexes is reduced by square root. It is natural to reduce RoPE frequencies in Eq. 9 by square root as
Note that for vision is often larger than that of language, and the number of frequencies is halved to cover both (x, y) dimensions with as well. This axial frequency has been used in a few pioneer works to further improve the performance of a new ViT architecture.
2.2 Mixed learnable frequency.
To handle mixed frequencies, we propose to use a rotation matrix in Eq. 10 in mixed axis form as
By using two frequencies for each axis, RoPE enables to handle the diagonal axis. The RoPE attention matrix in Eq. 8 is changed by mixed frequency as
This formulation is identical to the axial frequency implementation as or goes to zero. Thus, mixed frequency RoPE is a generalized version of axial frequency RoPE. Different from fixed frequencies in language RoPE and axial frequency, we let the network learn frequencies for as learnable parameters. Our mixed learnable frequency implementation enables diagonal direction handling to RoPE and makes RoPE learnable, like conventional positional embedding in the vision domain.Likeo RPB, we use separate sets of learnable frequencies for each head and every self-attention layer. It produces learnable parameters per self-attention layer. However, it is negligible since it requires only 0.01% of network parameters in ViT-B.
3 Discussion
This section introduces the interesting features of RoPE for ViT, offering insights into its functionality from different perspectives.
Vision models use diverse image resolutions depending on the goal of target tasks. For example, image classification uses as the standard resolution for comparison but utilizes small resolutions for training efficiency and enlarges resolutions to boost the performance additionally. Furthermore, object detection and segmentation prefer larger resolutions to capture small objects. Thus, transformers for vision should support resolution changes, which is linked to the necessity of resolution change in position embedding.
3.2 Phase shift in RoPE.
In sinusoidal representation, phase shift such as in is an important ability to control activation area. This phase shift ability is already included in and of the self-attention layer. Based on Eq. 8, when we apply and , the equation is
Thus, RoPE does not need additional parameters for phase shift since learnable parameters and can do the same role in network training.
3.3 Analyzing attention.
We analyze the attention matrix of RoPE ViT-B compared to the original ViT with APE. Following attention analysis in previous literature , we measure attention distances and entropy on ImageNet-1k validation set with various resolutions. Attention distance refers to the average spatial distance involved in attention interaction. The spatial distance between query-key is weighted and summed according to the attention probabilities. Note that we set the distance from the spatial token to the class token to zero for simplicity. Attention entropy represents the entropy values of attention probabilities, indicating probabilistic sharpness of attention. The measure is conducted for every self-attention layer except the last one since the last attention is only used for a class token. We use weights pretrained on ImageNet-1k in §4.1 and report average values across all tokens, attention heads, and validation samples.
The averaged attention distances for various resolutions are shown in Fig. 1. In training resolution , RoPE-Mixed increases attention distance at the middle layers but decreases it in the second and later layers. In other resolutions, the pattern is similar, but the difference is more significant than the training resolution. In short, RoPE-Mixed increases attention distance at the middle layers compared to APE, which becomes substantial at resolution changes. The entropy results are reported in Fig. 2 on multiple resolutions. Interestingly, the pattern is similar to attention distance. The entropy of RoPE-Mixed is larger than APE at middle layers. These analysis results imply that RoPE-Mixed makes attention interact with long-range (attention distance) and various tokens (entropy). We speculate that these differences in attention contributed to the performance improvement of RoPE observed in §4.
3.4 Computation costs.
Although RoPE has an involved formulation compared with APE and RPB, its computation cost is negligible to the overall computation. The rotation matrix in Eq. 12 and 14 is pre-computed before inference. The Hadamard product in Eq. 11 is the only computation required for inference - 1.8M FLOPs for ViT-B and accounts for only 0.01% of ViT-B’s 17.6G FLOPs.
Experiments
We apply 2D RoPE to two representative vision transformer architectures: ViT and Swin Transformer . Note that ViT uses APE, whereas Swin Transformer uses RPB. Thus, our experiment can verify the performance of RoPE when it replaces APE or RPB. RoPE in ViT and Swin Transformer is validated for vision recognition tasks, including multi-resolution classification (§4.1) on ImageNet-1k , object detection (§4.2) on MS-COCO , and semantic segmentation (§4.3) on ADE20k . On various tasks, we compare the conventional position embeddings (APE, RPB) with two variants of 2D RoPE RoPE-Axial (Eq. 12) and RoPE-Mixed (Eq. 14). Our experiments will exhibit the remarkable performance of 2D RoPE across all tasks, particularly with a significant margin in extrapolation.
Robustness on multi-resolution inputs is an essential factor of ViT performance, as it is closely related to their downstream performance in dense prediction tasks. In language models , RoPE exhibited strong extrapolation performance, i.e., text sequence longer than training samples. 2D RoPE might be suitable for large-resolution images, leveraging its extrapolation capabilities. We train ViTs and Swin Transformers on ImageNet-1k training set with high-performance training recipes . We report the accuracy on the ImageNet-1k validation set as varying image sizes. Note that we use the ImageNet-1k standard image resolution for training. Thus, a resolution larger than 224 is considered as extrapolation.
We apply 2D RoPE to ViT-S, ViT-B, and ViT-L. We train ViT with a strong supervised learning training recipe for ImageNet-1k, DeiT-III 400 epochs training recipe. When applying RoPE to ViT, we remove APE from ViT by default. Thus, 2D RoPE is the only position embedding for RoPE ViT. We denote ViT uses both RoPE and APE as RoPE+APE.
In Fig. 3, we compare 2D RoPE variants with APE for ViT position embedding. Both 2D RoPE, RoPE-Axial, and RoPE-Mixed implementations outperform APE for resolutions larger than 224, i.e., extrapolation cases. As expected, the strong extrapolation performance of RoPE can be extended to image recognition tasks. In comparison between RoPE-Axial and RoPE-Mixed, RoPE-Mixed performs better than RoPE-Axial in all input resolutions, meaning learnable frequencies for mixed axes are beneficial for classification.
We measure the performance of RoPE-Mixed when it is used with APE. The left side of Fig. 5 shows the performance of RoPE-Mixed with APE (RoPE-Mixed + APE) compared to RoPE-Mixed and APE. Note that we report accuracy improvement over APE for RoPE models to improve visualization. When used with RoPE, APE is beneficial for interpolation (res ) but reduces improvement on extrapolation (res ). RoPE+APE is almost double the improvement of RoPE-Mixed in interpolation, while the disadvantage in extrapolation is comparably small. Thus, RoPE+APE is a considerable choice for applying RoPE to ViT-based architectures on the target resolution of the tasks.
1.2 Swin Transformer
2D RoPE variants are applied to Swin Transformers, a milestone work in hierarchical ViT with relative position embedding RPB. The experiment in Swin Transformer investigates whether RoPE can replace RPB or work efficiently in a hierarchical ViT. We train Swin-T, Swin-S, and Swin-B on ImageNet-1k with 300 epochs of Swin Transformer training recipe . Similar to ViT, we replace RPB with 2D RoPE for comparison. Thus, RoPE Swin (i.e. Swin Transformer armed with RoPE) does not use RPB by default. A Swin Transformer using both position embedding is dubbed RoPE+RPE.
Fig 4 shows the multi-resolution performance of various Swin Transformers with different position embeddings. Two variants of 2D RoPE show remarkable performance improvements for extrapolation cases (res ). Even in interpolation (res ), RoPE-Mixed outperforms RPB by a large margin. It means that RoPE-Mixed is a more suitable option than RPB for Swin Transformers. When comparing RoPE-Mixed with RoPE-Axial, RoPE-Mixed outperforms in most resolutions. RoPE-Axial is especially weak in interpolation and significant extrapolation (res) cases.
We also measure performance when RoPE-Mixed is used together with RPB. The right side of Fig. 5 shows the results. Different from RoPE+APE in ViT, RoPE+RPB has no performance advantage compared to RoPE-Mixed in all resolutions. This implies that RoPE-Mixed effectively replaces RPB as a relative position embedding. Note that the gap between RoPE-Mixed and RoPE+RPB is significant when input resolution is far different from training resolution, demonstrating the advantage of RoPE-Mixed on resolution changes.
2 Object detection
We verify 2D RoPE in an object detection on MS-COCO dataset . DINO detector is trained using ViT and Swin Transformer as backbone network. We use ImageNet-1k supervised learning from §4.1 for backbone pre-training, and RoPE is only applied to the backbone network. The Detrex codebase is used for detection training. For ViT, we use the DINO-ViTDet 12 epochs training setting, and the DINO-Swin 12 epochs setting is selected for Swin-based DINO.
Table 1 shows the DINO-ViTDet results in bounding box AP. We report three variants of RoPEs: RoPE-Axial, RoPE-Mixed, and RoPE-Mixed + APE; all demonstrate remarkable performance improvements. DINO ViTDet achieves AP improvement of more than +1.0pp by changing positional embedding to RoPE. Among RoPE variants, RoPE-Mixed shows the best improvement at +1.8pp. AP in ViT-B and ViT-L. DINO-ViTDet uses ViT backbone with window-block attention, but still, a few layers remain as global attention. We believe that RoPE is highly effective due to the extrapolation in global attention, where it outperforms other position embeddings.
The performance of DINO-Swin is reported in Table 2. Like DINO-ViTDet, three variants, RoPE-Axial, RoPE-Mixed, and RoPE-Mixed + RPB, are reported. RoPE outperforms RPB for all variants. RoPE-Mixed performs better than RoPE-Axial, and +RPB is only beneficial in Swin-S. Performance improvement is smaller than DINO-ViTDet since DINO-Swin maintains a window size of the pre-trained backbone, i.e., DINO-Swin has no extrapolation. However, RoPE achieves meaningful gains and it has room for improvement by increasing the Swin Transformer’s window size for the detection backbone.
3 Semantic segmentation
We train 2D RoPE ViT and Swin Transformer for semantic segmentation tasks on ADE20k dataset. For ViT, we use UperNet with ViT backbone training recipe . For Swin, Mask2Former for semantic segmentation is used with the Swin backbone. ImageNet-1k supervised weights from §4.1 are used for backbone weights, the same as the detection setting. Also, RoPE is only applied to the backbone network. ViT-UperNet and Swin-Mask2Former are trained for 160k iterations.
Table 3 shows ViT-UperNet performance with APE and RoPE. RoPE-based models achieve impressive performance improvement in all cases. It is noteworthy that RoPE-Mixed + APE achieves +2.3 and +2.5 mIoU improvement with only position embedding changes. The improvement might originate from the extrapolation performance of RoPE since the ViT-UperNet setting uses images for backbone inputs. Among the three variants of 2D RoPE, RoPE-mixed + APE shows the best performance in all cases, which is different from detection results. As shown in Fig. 5, RoPE-mixed + APE has an advantage at interpolation while degrading performance at extrapolation. These results imply that the use of APE in a RoPE-based network should change depending on the target task. Swin-Mask2Former performances are shown in Table 4. RoPE also improves the performance of Swin-based segmentation. RoPE-Mixed shows the best performance among the three 2D RoPE variants. RoPE-Mixed + RPB has failed to improve RoPE Swin’s performance.
4 Comparison with multi-resolution methods
We compare 2D RoPE variants with recent ViT architecture designed for multi-resolution inference, namely ResFormer . ResFormer uses depth-wise convolutions as the position embedding method. It uses sinusoidal APE in Eq. 2 and depth-wise convolution after the patch-embed layer as Global Position Embedding (GPE). Also, another depth-wise convolution is used similar to skip-connection for every self-attention layer to add position embed as Local Position Embed (LPE). Using GPE and LPE, ResFormer is proposed as an improved ViT for multi-resolution inference. ResFormer is trained with multi-resolution training utilizing self-distillation loss. Since self-distillation with multi-resolution training is not a common recipe in ViT, we use ResFormer-S trained with fixed resolution and compare it with RoPE-Mixed ViT-S in §4.1. Table 5 shows a multi-resolution comparison of RoPE-Mixed with ResFormer-S-224. RoPE-Mixed outperforms ResFormer with a meaningful margin for extrapolation ranges (res ), but RoPE-Mixed shows performance lower than ResFormer for significant interpolation ranges (res ). To achieve comparable interpolation, RoPE-Mixed needs additional APE. Overall, the results show that RoPE-Mixed+APE outperforms ResFormer-S in multi-resolution inference.
Conclusion
Rotary Position Embedding (RoPE) is a novel method for relative position embedding with a lot of potential. However, it has been underexplored in vision modeling. In this paper, we have conducted a comprehensive investigation of 2D RoPE for Vision Transformer (ViT) and proposed an improved 2D RoPE, RoPE-Mixed, utilizing mixed axis frequency with learnable parameters. Our experiments show that 2D RoPE is an effective solution for multi-resolution classification for both ViT and Swin Transformers, particularly for large resolutions. 2D RoPE shows improved performance with a significant margin in downstream tasks, such as object detection and semantic segmentation. It is noteworthy that our RoPE-Mixed outperforms conventional 2D RoPE in various tasks, further enhancing the contribution of this research. We believe that our study will be useful for vision researchers looking for state-of-the-art performance by suggesting 2D RoPE as a solution for them.
References
Appendix
Appendix 0.A Experiments (cont’d)
We demonstrated the performance of 2D RoPE with performance graphs through various input resolutions in the main paper. This appendix provides the entire performance numbers for the multi-resolution experiments. We measured the performance of ViTs and Swin Transformers with default position embedding (APE or RPB), 2D RoPE variants (RoPE-Axial and RoPE-Mixed), and 2D RoPE variants with default position embedding (RoPE-Axial and RoPE-Mixed + APE or RPB). Note that figures in the paper do not include small resolutions such as to improve the visualization.
Table 0.A.1, 0.A.2, and 0.A.3 report the total numbers of multi-resolution classification, which is illustrated in Figure 3. 2D RoPE variants outperform APE in the smallest resolution with significant gap.
A.2 Multi-resolution classification – Swin Transformer
Table 0.A.4, 0.A.5, and 0.A.6 show the total numbers of multi-resolution classification of Swin Transformer with 2D RoPE variants corresponding to Figure 4 of paper. Similar to ViT cases, 2D RoPE variants significantly outperform RPB in small resolutions: and .
A.3 2D RoPE with APE or RPB
In Figure 5 of the paper, we report the performance improvement of RoPE-Mixed compared to base position embeddings: APE or RPB. We provide numbers for Figure 5 in Table 0.A.7 and 0.A.8. Note that each number means performance improvement (%p.) compared to base position embeddings (APE or RPB).