UniFormer: Unifying Convolution and Self-attention for Visual Recognition

Kunchang Li, Yali Wang, Junhao Zhang, Peng Gao, Guanglu Song, Yu Liu, Hongsheng Li, Yu Qiao

Introduction

Representation learning is a fundamental research topic for visual recognition . Basically, we confront two distinct challenges that exist in visual data such as images and videos. On one hand, the local redundancy is large, e.g., visual content in a local region (space, time or space-time) tends to be similar. Such locality often introduces inefficient computation. On the other hand, the global dependency is complex, e.g., targets in different regions have dynamic relations. Such long-range interaction often causes ineffective learning.

To tackle such difficulties, researchers have proposed a number of powerful models in visual recognition . In particular, the mainstream backbones are Convolution Neural Networks (CNNs) and Vision Transformers (ViTs) , where convolution and self-attention are the key operations in these two structures. Unfortunately, each of these operations mainly addresses one aforementioned challenge while ignoring the other. For example, the convolution operation is good at reducing local redundancy and avoiding unnecessary computation, by aggregating each pixel with context from a small neighborhood (e.g, 3×\times3 or 3×\times3×\times3). However, the limited receptive field makes convolution suffer from difficulty in learning global dependency . Alternatively, self-attention has been recently highlighted in the ViTs. By similarity comparison among visual tokens, it exhibits the strong capacity of learning global dependency in both images and videos . Nevertheless, we observe that ViTs are often inefficient to encode local features in the shallow layers.

We take the well-known ViTs in the image and video domains (i.e., DeiT and TimeSformer ) as examples, and visualize their attention maps in the shallow layer. As shown in Figure 2, both ViTs indeed capture detailed visual features in the shallow layer, while spatial and temporal attention are redundant. One can easily see that, given an anchor token, spatial attention largely concentrates on the tokens in the local region (mostly 3×\times3), and learns little from the rest tokens in this image. Similarly, temporal attention mainly aggregates the tokens in the adjacent frames, while losing sight of the rest tokens in the distant frames. However, such local focus is obtained by global comparison among all the tokens in space and time. Clearly, this redundant attention manner brings large and unnecessary computation burden, thus deteriorating the computation-accuracy balance in ViTs (Figure 1).

Based on these discussions, we propose a novel Unified transFormer (UniFormer) in this work. It flexibly unifies convolution and self-attention in a concise transformer format, which can tackle both local redundancy and global dependency for effective and efficient visual recognition. Specifically, our UniFormer block consists of three key modules, i.e., Dynamic Position Embedding (DPE), Multi-Head Relation Aggregator (MHRA), and Feed-Forward Network (FFN). The distinct design of the relation aggregator is the key difference between our UniFormer and the previous CNNs and ViTs. In the shallow layers, our relation aggregator captures local token affinity with a small learnable parameter matrix, which inherits the convolution style that can largely reduce computation redundancy by context aggregation in the local region. In the deep layers, our relation aggregator learns global token affinity with token similarity comparison, which inherits the self-attention style that can adaptively build long-range dependency from distant regions or frames. Via progressively stacking local and global UniFormer blocks in a hierarchical manner, we can flexibly integrate their cooperative power to promote representation learning. Finally, we provide a generic and powerful backbone for visual recognition and successfully address various downstream vision tasks with simple and elaborate adaptations. Additionally, we further introduce the lightweight design for UniFormer, which can achieve a preferable accuracy-throughout balance, by a concise hourglass style of token shrinking and recovering.

Extensive experiments demonstrate the strong performance of our UniFormer on a broad range of vision tasks, including image classification, video classification, object detection, instance segmentation, semantic segmentation and pose estimation. Without any extra training data, UniFomrer-L achieves 86.3 top-1 accuracy on ImageNet-1K. Moreover, with only ImageNet-1K pre-training, UniFormer-B achieves 82.9/84.8 top-1 accuracy on Kinetics-400/Kinetics-600, 60.9 and 71.2 top-1 accuracy on Something-Something V1&V2, 53.8 box AP and 46.4 mask AP on the COCO detection task, 50.8 mIoU on the ADE20K semantic segmentation task, and 77.4 AP on the COCO pose estimation task. Finally, our efficient UniFormer with a concise hourglass design can achieve 2-4×\bm{\times} higher throughput than the recent lightweight models.

Related Work

In the past few years, the development of computer vision has been mainly driven by convolutional neural networks (CNNs). Beginning with the classical AlexNet , many powerful CNN networks have been proposed and achieved remarkable performance in various tasks of image understanding . Recently, due to the fact that video has gradually become one main data resource in many realistic applications, researchers have attempted to apply CNNs in the video domain. Naturally, one can adapt 2D convolution as 3D one, by temporal dimension extension . However, 3D CNNs often suffer from difficult optimization problem and large computation cost. To resolve these issues, the prior works try to inflate the pre-trained 2D convolution kernels for better optimization and factorize 3D convolution kernels in different dimensions to reduce complexity . Besides, many recent studies of video understanding focus on adapting vanilla 2D CNNs with elaborated temporal modeling modules, such as temporal shift , motion enhancement , and spatiotemporal excitation , etc. Unfortunately, due to the limited reception field, the traditional convolution struggles to capture long-range dependency even if they are stacked deeper.

2 Vision Transformers (ViTs)

To capture long-term dependencies, Vision Transformer (ViT) has been proposed . With the inspiration of Transformer architectures in NLP , ViT treats image as a number of visual tokens and leverages attention to encode token relations for representation learning. However, vanilla ViT depends on sufficient training data and careful data augmentation. To tackle these problems, several approaches have been developed by improved patch embedding , data-efficient training , efficient self-attention , and multi-scale architectures . These works successfully boost the performance of ViT on various image tasks . Recently, researchers have attempted to extend image ViTs for video modeling. The classical work is TimeSformer by spatial-temporal attention. Starting from this, many works propose different variants for spatiotemporal representation learning , and subsequently they are adapted to various video understanding tasks . Although these works demonstrate the outstanding ability of ViTs to learn long-term token relations, the self-attention mechanism requires costly token-to-token comparisons . Hence, it is often inefficient to encode low-level features, as shown in Figure 2. Though Video Swin advocates an inductive bias of locality with shift window, window-based self-attention is still less efficient than local convolution when encoding low-level features. Moreover, the shifted window should be carefully configured.

3 Combination of CNNs and ViTs

To bridge the gap between CNNs and ViTs, researchers have tried to take advantage of them to build stronger vision backbones for image understanding, by adding convolutional patch stem for fast convergence , introducing convolutional position embedding , inserting depthwise convolution into feed-forward network , utilizing convolutional projection in self-attention , and combining MBConv with Transformer . As for video understanding, the combination is also straightforward, i.e., one can insert self-attention as global attention , and/or use convolution as patch stem . However, all these approaches ignore inherent relations between convolution and self-attention, leading to inferior local and/or global token relation learning. Several recent works have demonstrated that self-attention operates similarly to convolution . But they suggest replacing convolution instead of combining them together. Differently, our UniFormer unifies convolution and self-attention in the transformer style, which can effectively learn local and global token relation, and achieve better accuracy-computation trade-offs on all the vision tasks from image to the video domain.

4 Lightweight CNNs and ViTs

In many practical applications, the running platforms usually lack enough computility. Hence, a series of lightweight CNNs are proposed to satisfy such on-device requirements. For example, the classical MobileNets adopt depthwise separable convolution in well-organized efficient ResNet. ShuffleNets leverage channel shuffle for computation reduction. EfficientNets further scale model with neural architecture search at depth, width, and resolution. However, such lightweight design has not been fully investigated in ViTs. Two recent works are proposed by designing transformers as convolutions (i.e., MobielViT), and introducing a parallel architecture of MobileNet and ViT (i.e., MobileFormer). But both works ignore the inference speed (e.g., throughout). Hence, the prior efficient CNNs are still the better choice. To bridge this gap, we build a lightweight UniFormer by token shrinking and recovering in Section 5.

Method

In this section, we introduce the proposed UniFormer in detail. First, we describe the overview of our UniFormer block. Then, we explain its key modules such as multi-head relation aggregator and dynamic position embedding. Moreover, we discuss the distinct relations between our UniFormer and existing CNNs/ViTs, showing its preferable design for accuracy-computation balance.

Figure 3 shows our Unified transFormer (UniFormer). For simple description, we take a video with TT frames as an example and an image input can be seen as a video with a single frame. Hence, the dimensions highlighted in red only exit for the video input, while all of them are equal to one for image input. Our UniFormer is a basic transformer format, while we elaborately design it to tackle computational redundancy and capture complex dependency.

Specifically, our UniFormer block consists of three key modules: Dynamic Position Embedding (DPE), Multi-Head Relation Aggregator (MHRA) and Feed-Forward Network (FFN):

2 Multi-Head Relation Aggregator

As analyzed before, the traditional CNNs and ViTs focus on addressing either local redundancy or global dependency, leading to unsatisfactory accuracy or/and unnecessary computation. To overcome these difficulties, we introduce a generic Relation Aggregator (RA), which elegantly unifies convolution and self-attention for token relation learning. It can achieve efficient and effective representation learning by designing local and global token affinity in the shallow and deep layers respectively. Specifically, MHRA exploits token relationships in a multi-head style:

As shown in Figure 2, though the previous ViTs compare similarities among all the tokens, they finally learn local representations. Such redundant self-attention design brings large computation cost in the shallow layers. Based on it, we suggest learning token affinity in a small neighborhood, which coincidentally shares a similar insight with the design of a convolution filter. Hence, we propose to represent local affinity as a learnable parameter matrix in the shallow layers. Concretely, given an anchor token Xi\mathbf{X}_{i}, our local RA learns the affinity between this token and other tokens in the small neighborhood Ωit×h×w\Omega_{i}^{t\times h\times w} (t<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>=</mo></mrow><annotationencoding="application/x−tex">=</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.3669em;"></span><spanclass="mrel">=</span></span></span></span></span>1t<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>=</mo></mrow><annotation encoding="application/x-tex">=</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.3669em;"></span><span class="mrel">=</span></span></span></span></span>1 for an image input):

Comparison to Convolution Block. Interestingly, we find that our local MHRA can be interpreted as a generic extension of MobileNet block . Firstly, the linear transformation V(⋅){\rm V}(\cdot) in Eq. 4 is equivalent to a pointwise convolution (PWConv), where each head is corresponding to an output feature channel Vn(X){\rm V}_{n}({\rm X}). Furthermore, our local token affinity Anlocal{\rm A}_{n}^{local} can be instantiated as the parameter matrix that operated on each output channel (or head) Vn(X){\rm V}_{n}(\mathbf{X}), thus the relation aggregator Rn(X)=AnlocalVn(X){\rm R}_{n}(\mathbf{X})={\rm A}_{n}^{local}{\rm V}_{n}(\mathbf{X}) can be explained as a depthwise convolution (DWConv). Finally, the linear matrix U\mathbf{U}, which concatenates and fuses all heads, can also be seen as a pointwise convolution. As a result, such local MHRA can be reformulated with a manner of PWConv-DWConv-PWConv in the MobileNet block. In our experiments, we instantiate our local MHRA as such channel-separated convolution, so that our UniFormer can boost computation efficiency for visual recognition. Moreover, different from the MobileNet block, our local UniFormer block is designed as a generic transformer format, i.e., it also contains dynamical position encoding (DPE) and feed-forward network (FFN), besides MHRA. This unique integration can effectively enhance token representation, which has not been explored in the previous convolution blocks.

2.2 Global MHRA

In the deep layers, it is important to exploit long-range relation in the broader token space, which naturally shares a similar insight with the design of self-attention. Therefore, we design the token affinity via comparing content similarity among all the tokens:

where Xj\mathbf{X}_{j} can be any token in the global tube with a size of T<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>H<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>WT<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>H<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>W (T<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>=</mo></mrow><annotationencoding="application/x−tex">=</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.3669em;"></span><spanclass="mrel">=</span></span></span></span></span>1T<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>=</mo></mrow><annotation encoding="application/x-tex">=</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.3669em;"></span><span class="mrel">=</span></span></span></span></span>1 for an image input), while Qn(⋅)Q_{n}(\cdot) and Kn(⋅)K_{n}(\cdot) are two different linear transformations.

Comparison to Transformer Block. Our global MHRA Anglobal{\rm A}_{n}^{global} (Eq. 7) can be instantiated as a spatiotemporal self attention, where Qn(⋅){\rm Q}_{n}(\cdot), Kn(⋅){\rm K}_{n}(\cdot) and Vn(⋅){\rm V}_{n}(\cdot) become Query, Key and Value in ViT . Hence, it can effectively learn long-range dependency. However, our global UniFormer block is different from the previous ViT blocks. First, most video transformers divide spatial and temporal attention in the video domain , in order to reduce the dot-product computation in token similarity comparison. But such an operation inevitably deteriorates the spatiotemporal relation among tokens. In contrast, our global UniFormer block jointly encodes spatiotemporal token relation to generate more discriminative video representation for recognition. Since our local UniFormer block largely saves computation of token comparison in the shallow layers, the overall model can achieve a preferable computation-accuracy balance. Second, instead of absolute position embedding , we adopt dynamic position embedding (DPE) in our UniFormer. It is in convolution style (see the next section), which can overcome permutation-invariance and be friendly to different input lengths of visual tokens.

3 Dynamic Position Embedding

The position information is an important clue to describe visual representation. Previously, most ViTs encode such information by absolute or relative position embedding . However, absolute position embedding has to be interpolated for various input sizes with fine-tuning , while relative position embedding does not work well due to the modification of self-attention . To improve flexibility, convolutional position embedding has been recently proposed . In particular, conditional position encoding (CPE) can implicitly encode position information via convolution operators, which unlocks Transformer to process arbitrary input size and promotes recognition performance. Due to its plug-and-play property, we flexibly adopt it as our Dynamical Position Embedding (DPE) in the UniFormer:

where DWConv{\rm DWConv} refers to depthwise convolution with zero paddings. We choose such a design as our DPE based on the following reasons. First, depthwise convolution is friendly to arbitrary input shapes, e.g., it is straightforward to use its spatiotemporal version to encode 3D position information in videos. Second, depthwise convolution is light-weight, which is an important factor for computation-accuracy balance. Finally, we add extra zero paddings, since it can help tokens be aware of their absolute positions by querying their neighbors progressively .

Framework

In the section, we mainly develop visual frameworks for various downstream tasks. Specifically, we first develop a number of visual backbones for image classification, by hierarchically stacking our local and global UniFormer blocks with consideration of computation-accuracy balance. Then, we extend the above backbones to tackle other representative vision tasks, including video classification and dense prediction (i.e., object detection, semantic segmentation and human pose estimation). Such generality and flexibility of our UniFormer demonstrate its valuable potential for computer vision research and beyond.

It is important to progressively learn visual representation for capturing semantics in the image. Hence, we build up our backbone with four stages, as illustrated in Figure 3.

More specifically, we use the local UniFormer blocks in the first two stages to reduce computation redundancy, while the global UniFormer blocks are utilized in the last two stages to learn long-range token dependency. For the local UniFormer block, MHRA is instantiated as PWConv-DWConv-PWConv with local token affinity (Eq. 6), where the spatial size of DWConv is set to 5×\times5 for image classification. For the global UniFormer block, MHRA is instantiated as multi-head self-attention with global token affinity (Eq. 7), where the number of attention heads is set to 64. For both local and global UniFormer blocks, DPE is instantiated as DWConv with a spatial size of 3×\times3, and the expand ratio of FFN is 4.

Additionally, as suggested in the CNN and ViT literatures , we utilize BN for convolution and LN for self-attention. For feature downsampling, we use the 4×\times4 convolution with stride 4×\times4 before the first stage and the 2×\times2 convolution with stride 2×\times2 before other stages. Besides, an extra LN is added after each downsampling convolution. Finally, the global average pooling and fully connected layer are applied to output the predictions. When training models with Token Labeling , we add another fully connected layer for auxiliary loss. For various computation requirements, we design three model variants as shown in Table I.

2 Video Classification

Given our image-based 2D backbones, one can easily adapt them as 3D backbones for video classification. Without loss of generality, we adjust Small and Base models for spatiotemporal modeling. Specifically, the model architectures keep the same with four stages, where we use the local UniFormer blocks in the first two stages and the global UniFormer blocks in the last two stages. But differently, all the 2D convolution filters are changed as 3D ones via filter inflation . Concretely, the kernel size of DWConv in DPE and local MHRA are 3×\times3×\times3 and 5×\times5×\times5 respectively. Moreover, we downsample both spatial and temporal dimensions before the first stage. Hence, the convolution filter before this stage becomes 3×\times4×\times4 with the stride of 2×\times4×\times4. For the other stages, we just downsample the spatial dimension to decrease the computation cost and maintain high performance. Hence, the convolution filters before these stages are 1×\times2×\times2 with stride of 1×\times2×\times2 .

Note that we use spatiotemporal attention in the global UniFormer blocks for learning token relation jointly in the 3D view. It is worth mentioning that, due to the large model sizes, the previous video transformers divide spatial and temporal attention to reduce computation and alleviate overfitting, but such factorization operation inevitably tears spatiotemporal token relations. In contrast, our joint spatiotemporal attention can avoid the issue. Besides, our local UniFormer blocks largely save computation via 3D DWconv. Hence, our model can achieve effective and efficient video representation learning.

3 Dense Prediction

Dense prediction tasks are necessary to verify the generality of our recognition backbones. Hence, we adopt our UniFormer backbones for a number of popular dense tasks such as object detection, instance segmentation, semantic segmentation, and human pose estimation. However, direct usage of our backbone is not suitable because of the high input resolution of most dense prediction tasks, e.g., the size of the image is 1333×\times800 in the COCO object detection dataset. Naturally, feeding such images into our classification backbones would inevitably lead to large computation, especially when operating self-attention of global UniFormer block in the last two stages. Taking h<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>wh<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>w visual tokens as an example, the MatMul operation in token similarity comparison (Eq. 7) causes O(w2h2){\mathcal{O}}(w^{2}h^{2}) complexity, which is prohibitive for most dense tasks.

We propose to adjust the global UniFormer block for different downstream tasks. First, we analyze the FLOPs of our UniFormer-S under different input resolutions. Figure 4 clearly indicates that Relation Aggregator (RA) in Stage3 occupies large computation. For example, for a 1008×\times1008 image, the MatMul operation of RA in Stage3 even occupies over 50% of the total FLOPs, while the FLOPs in Stage4 is only 1//28 of that in Stage3. Thus we focus on modifying RA in Stage3 for computation reduction.

Inspired by , we propose to apply our global MHRA in a predefined window (e.g., 14×\times14), instead of using it in the entire image with high resolution. Such operation can effectively cut down computation with the complexity of O(whp2){\mathcal{O}}(whp^{2}), where pp is the window size. However, it undoubtedly drops model performance, due to insufficient token interaction. To bridge this gap, we integrate window and global UniFormer blocks together in Stage3, where a hybrid group consists of three window blocks and one global block. In this case, there are 2/5 hybrid groups in Stage3 of our UniFormer-Small/Base backbones.

Based on this design, we next introduce the specific backbone settings of various dense tasks, depending on the input resolution of training and testing images. For object detection and instance segmentation, the input images are usually large (e.g., 1333×\times800), thus we adopt the hybrid block style in Stage3. In contrast, the inputs are relatively small for pose estimation, such as 384×\times288, hence global blocks are still applied in Stage3 for both training and testing. Specially, for semantic segmentation, the testing images are often larger than the training ones. Therefore, we utilize the global blocks in Stage3 for training, while adapting the hybrid blocks in Stage3 for testing. We use the following simple design. The window-based block in testing has the same receptive field as the global block in training, e.g., 32×\times32. Such design can maintain training efficiency, and boost testing performance by keeping consistency with training as much as possible.

Towards Lightweight UniFormer

Recently, researchers have tried to combine CNNs with ViTs to design lightweight models. For example, MobileFormer proposes a parallel design of MobileNet and ViT, and MobileViT designs transformers as convolutions. However, the inference speeds of these works should be further improved. Hence, the prior efficient CNNs are still the better choice, such as EfficientNet for image tasks and MoViNet for video tasks. To bridge the gap, we propose a lightweight UniFormer architecture in Figure 6, by designing a distinct hourglass UniFormer block. In this block, we adaptively leverage token shrinking and recovering, to achieve a preferable accuracy-throughput balance.

Since a large computation load lies in token similarity comparison in the global UniFormer block, we propose an Hourglass UniFormer (H-UniFormer) block to reduce the number of visual tokens involved in global MHRA. Note that, the existing token pruning methods are infeasible for our UniFormer and those ViTs with convolution . The main reason is that, after pruning, the rest tokens often maintain a broken spatiotemporal structure, which makes convolution inapplicable. To overcome such difficulty, we propose a concise integration of token shrinking and recovering in our H-UniFormer block.

Token Recovering. After learning token interactions in global MHRA and FFN, we replicate the representative token to recover the unimportant tokens. Thus, we can maintain the spatiotemporal structure of all the visual tokens, for effective dynamic position encoding (i.e., convolution) in the next H-UniFormer block.

2 LightWeight UniFormer Architecture

To build our light-weight UniFormer, We follow most of the architectures in Section 4.1, except that we change to use the H-UniFormer block in Stage3 and Stage4, and adopt smaller depth, width or resolution (e.g., 128). Specifically, for UniFomrer-XS, the block number, channel number and head dimension are , and 32. For UniFomrer-XXS, the block number, channel number and head dimension are , and 28.

Beginning from the second layer in Stage3, we utilize the similarity score in the previous layer Apre\mathbf{A}^{pre} to guide the token shrinking. Based on the phenomenon that the locations of the crucial tokens are basically the same among different layers , we update the similarity scores via mean, i.e., A=(A+Apre)/2\mathbf{A}=(\mathbf{A}+\mathbf{A}^{pre})/2. Thus it can focus on the significant tokens consistently. By default, we keep half of the tokens and fuse the rest (i.e., the shrinking ratio is 0.5) in our light-weight UniFormer.

Experiments

To verify the effectiveness and efficiency of our UniFormer for visual recognition, we conduct extensive experiments on ImageNet-1K image classification, Kinetics-400/600 and Something-Something V1&V2 video classification, COCO object detection, instance segmentation and pose estimation, and ADE20K semantic segmentation. We also perform comprehensive ablation studies to analyze each design of our UniFormer.

Settings. We train our models from scratch on the ImageNet-1K dataset . For a fair comparison, we follow the same training strategy proposed in DeiT by default, including strong data augmentation and regularization. Additionally, we set the stochastic depth rate as 0.1/0.3/0.4 respectively for our UniFormer-S/B/L in Table I. We train all models via AdamW optimizer with cosine learning rate schedule for 300 epochs, while the first 5 epochs are utilized for linear warm-up . The weight decay, learning rate and batch size are set to 0.05, 1e-3 and 1024 respectively. For UniFormer-S†\dagger, we follow state-of-the-art ViTs to apply overlapped patch embedding and blocks (3/5/9/3 blocks in each stage) for fair comparisons. As for UniFormer-B, we use the learning rate of 8e-4 for better convergence.

For training high-performance ViTs, hard distillation and Token Labeling are proposed, both of which are complementary to our backbones. Since Token Labeling is more efficient, we apply it with an extra fully connected layer and auxiliary loss, following the settings in LV-ViT . Different from the training settings in DeiT, MixUp and CutMix are not used since they conflict with MixToken . The base learning rate is 1.6e-3 for the batch size of 1024 by default. Specially, we adopt the base learning rate of 1.2e-3 and layer scale for UniFormer-L to avoid NaN loss. When fine-tuning our models on larger resolution, i.e., 384×\times384, the weight decay, learning rate, batch size, warm-up epoch and total epoch are set to 1e-8, 5e-6, 512, 5 and 30.

Results. In Table II, we compare our UniFormer with the state-of-the-art CNNs, ViTs and their combinations. It clearly shows that our UniFormer outperforms previous models under different computation restrictions. For example, our UniFormer-S†\dagger achieves 83.4% top-1 accuracy with only 4.2G FLOPs, surpassing RegNetY-4G , Swin-T , CSwin-T and CoAtNet by 3.4%, 2.1%, 0.7% and 1.8% respectively. Though EfficientNet comes from extensive neural architecture search, our UniFormer outperforms it (83.9% vs. 83.6%) with less computation cost (8.3G vs. 9.9G). Furthermore, we enhance our models with Token Labeling , which is denoted by ‘⋆’. Compared with the models training with the same settings, our UniFormer-L achieves higher accuracy but only 21% FLOPs of LV-ViT-M and 61% FLOPs of VOLO-D3. Moreover, when fine-tuned on 384×\times384 images, our UniFormer-L obtains 86.3% top-1 accuracy. It is even better than EfficientNetV2-L with larger input, demonstrating the powerful learning capacity of our UniFormer.

2 Video Classification

Settings. We evaluate our UniFormer on the popular Kinetics-400 , Kinetics-600 , UCF101 and HMDB51 , and we verify the transfer learning performance on temporal-related datasets Something-Something (SthSth) V1&V2 . Our codes mainly rely on PySlowFast . For training, we adopt the same training strategy as MViT by default, but we do not apply random horizontal flip for SthSth. We utilize AdamW optimizer with cosine learning rate schedule to train our video backbones.. The first 5 or 10 epochs are used for warm-up to overcome early optimization difficulty. For UniFormer-S, the warmup epoch, total epoch, stochastic depth rate, weight decay are set to 10, 110, 0.1 and 0.05 respectively for Kinetics, 5, 50, 0.3 and 0.05 respectively for SthSth, and 5, 20, 0.2 and 0.05 for UCF101 and HMDB51. For UniFormer-B, all the hyper-parameters are the same unless the stochastic depth rates are doubled. Moreover, We linearly scale the base learning rates according to the batch size, which are 1e-4⋅batchsize32\cdot\frac{batchsize}{32} for Kinetics, 2e-4⋅batchsize32\cdot\frac{batchsize}{32} for SthSth, and 1e-5⋅batchsize32\cdot\frac{batchsize}{32} for UCF101 and HMDB51.

We utilize the dense sampling strategy for Kinetics, UCF101 and HMDB51, and uniform sampling strategy for Something-Something. To reduce the total training cost, we inflate the 2D convolution kernels pre-trained on ImageNet for Kinetics . To obtain a better FLOPs-accuracy balance, Besides, we adopt multi-clip testing for Kinetics and multi-crop testing for Something-Something. All scores are averaged for the final prediction.

Results on Kinetics. In Table III, we compare our UniFormer with the state-of-the-art methods on Kinetics-400 and Kinetics-600. The first part shows the prior works using CNN. Compared with SlowFast equipped with non-local blocks , our UniFormer-S16f requires 42×\mathbf{42}\times fewer GFLOPs but obtains 1.0% performance gain on both datasets (80.8% vs. 79.8% and 82.8% vs. 81.8%). Even compared with MoViNet , which is a strong CNN-based models via extensive neural architecture search, our model achieves slightly better results (82.0% vs. 81.5%) with fewer input frames (16f<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>416f<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>4 vs. 120f120f). The second part lists the recent methods based on vision transformers. With only ImageNet-1K pre-training, UniFormer-B16f surpasses most existing backbones with large dataset pre-training. For example, compared with ViViT-L pre-trained from JFT-300M and Swin-B pre-trained from ImageNet-21K, UniFormer-B32f obtains comparable performance (82.9% vs. 82.8% and 82.7%) with 16.7×\mathbf{16.7\times} and 3.3×\mathbf{3.3\times} fewer computation on both Kinetics-400 and Kinetics-600. These results demonstrate the effectiveness of our UniFormer for video.

Results on Something-Something. Table IV presents the results on Something-Something (SthSth) V1&V2. Since these datasets require robust temporal relation modeling, it is difficult for the CNN-based methods to capture long-term dependencies, which leads to their worse results. On the contrary, video transformers are good at processing long sequential data and demonstrate better transfer learning capabilities , thus they achieve higher accuracy but with large computation costs. In contrast, our UniFormer-S16f combines the advantages of both convolution and self-attention, obtaining 54.4%/65.0% in SthSth V1/V2 with only 42 GFLOPs. It also demonstrates that small model UniFormer-S benefit from larger dataset pre-training (Kinetics-400, 53.8% vs. Kinetics-600, 54.4%), but large model UniFormer-B do not (Kinetics-400, 59.1% vs. Kinetics-600, 58.8%). We argue that the large model is easy to converge better. Besides, it is worth noting that our UniFormer pre-trained from Kinetis-600 outperforms all the current methods under the same settings. In fact, our best model achieves the new state-of-the-art results: 61.0% top-1 accuracy on SthSth V1 (4.2% higher than TDNEN) and 71.2% top-1 accuracy on SthSth V2 (1.6% higher than Swin-B ). Such results verify its high capability of spatiotemporal learning.

Results on UCF101 and HMDB51. We further verify the generalization ability on UCF101 and HMDB51. Since these datasets are relatively small, the performances have already saturated. As shown in Table VI, our UniFormer significantly outperforms the previous SOTA methods, revealing its strong generalization ability to transfer to small datasets.

3 Object Detection and Instance Segmentation

Settings. We benchmark our models on object detection and instance segmentation with COCO2017 . The ImageNet-1K pre-trained models are utilized as backbones and then armed with two representative frameworks: Mask R-CNN and Cascade Mask R-CNN . Our codes are mainly based on mmdetection , and the training strategies are the same as Swin Transformer . We adopt two training schedules: 1×\times schedule with 12 epochs and 3×\times schedule with 36 epochs. For the 1×\times schedule, the shorter side of the image is resized to 800 while keeping the longer side no more than 1333. As for the 3×\times schedule, we apply the multi-scale training strategy to randomly resize the shorter side between 480 to 800. Besides, we use AdamW optimizer with the initial learning rate of 1e-4 and weight decay of 0.05. To regularize the training, we set the stochastic depth drop rates to 0.1/0.3 and 0.2/0.4 for our small/base models with Mask R-CNN and Cascade Mask R-CNN.

Results. Table V reports box mAP (APbAP^{b}) and mask mAP (APmAP^{m}) of the Mask R-CNN framework. It shows that our UniFormer variants outperform all the CNN and Transformer backbones. To reduce the training cost for object detection, we utilize a hybrid UniFomer style with a window size of 14 in Stage3 (denoted by h14h14). Specifically, with 1×\times schedule, our UniFormer brings 7.0-7.6 points of box mAP and 6.7-7.2 mask mAP against ResNet at comparable settings. Compared with the popular Swin Transformer , our UniFormer achieves 2.6-3.4 points of box mAP and 2.2-2.5 mask mAP improvement. Moreover, with 3×\times schedule and multi-scale training, our models still consistently surpass CNN and Transformer counterparts. For example, our UniFormer-B outperforms the powerful CSwin-S by +0.3 box mAP and +0.3 mask mAP, and even better than larger backbones such as Swin-B and Focal-B . Table VII reports the results with the Cascade Mask R-CNN framework. The consistent improvement demonstrates our stronger context modeling capacity.

4 Semantic Segmentation

Settings. Our semantic segmentation experiments are conducted on the ADE20k dataset and our codes are based on mmseg . We adopt the popular Semantic FPN and Upernet as the basic framework. For a fair comparison, we follow the same setting of PVT to train Semantic FPN for 80k iterations with cosine learning rate schedule . As for Upernet, we apply the settings of Swin Transformer with 160k iteration training. The stochastic depth drop rates are set to 0.1/0.2 and 0.25/0.4 for small/base variants with Semantic FPN and Upernet respectively.

Results. Table VIII and Table IX report the results of different frameworks. It shows that with the Semantic FPN framework, our UniFormer-Sh32/Bh32 achieve +4.7/+2.5 higher mIoU than the Swin Transformer with similar model sizes. When equipped with the UperNet framework, they achieve +2.5/+1.9 mIoU and +2.7/+1.2 MS mIoU improvement. Furthermore, when utilizing the global MHRA, the results are consistently improved but with a larger computation cost. More results can be found in Table XXII.

5 Pose Estimation

Settings. We evaluate the performance of UniFormer on the COCO2017 human pose estimation benchmark. For a fair comparison with previous SOTA methods, we employ a single Top-Down head after our backbones. We follow the same training and evaluation setting of mmpose as HRFormer . In addition, the batch size and stochastic depth drop rates are set to 1024/256 and 0.2/0.5 for small/base variants during training.

Results. Table X reports results of different input resolutions on COCO validation set. Compared with previous SOTA CNN models, our UniFormer-B surpasses HRNet-W48 by 0.4% AP with fewer parameters (53.5M vs. 63.6M) and FLOPs (22.1G vs. 32.9G). Moreover, our UniFormer-B can outperform the current best approach HRFormer by 0.2% AP with smaller FLOPs (29.6G vs. 30.7G). It is worth noting that HRFormer , PRTR , TokenPose and TransPose are sophisticatedly designed for pose estimation task. On the contrary, our UniFormer can outperform all of them as a simple yet effective backbone.

6 Light-Weight UniFormer

Settings. For the light-weight UniFormer, we follow most of the previous settings. Differently, we train UniFormer-XSS and UniFormer-XS for 600 epochs on ImageNet, since the lightweight models are difficult to converge

Results of classification. Table XI represents the results on ImageNet. We roughly divide the models according to the FLOPs: <<1G and 1G−-2G. It clearly reveals that our efficient UniFormer achieves the best accuracy-throughput trade-off under similar FLOPs. For example, compared with the strong CNN method EfficientNet-B3, our UniFormer-XS192 obtains 1.7×\times higher throughput with similar performance. Compared with SOTA MobileFormer that combines CNN and ViT, our UniFormer-XXS192 obtains 0.6% higher accuracy with 16% higher throughput. We further fine-tune the above models with different frames on Kinetics-400. Results in Table XII show that our efficient backbone surpasses the SOTA lightweight video backbones by a large margin. Compared with MoViNet-A0, our UniFormer-XXS150×16f achieves 9.3% higher performance with 16% higher throughput. While compared with X3D-S, our UniFormer-XS192×32f runs 4.2×\times faster with 5.4% higher accuracy. Note that we do not apply complicated designs as in the recent lightweight methods. Our concise extension already shows powerful performance, which further demonstrates the great potential of UniFormer.

Results of dense prediction. We also verify the efficient UniFormer for COCO object detection and instance segmentation in Table XIII, and ADE20K semantic segmentation in Table XIV. Our models obviously beat ResNet and PVTv2 on these dense prediction tasks. For example, our UniFormer-XS brings 2.7 points of box mAP and 2.1 mask mAP against PVTv2-B1 on COCO, and achieves +1.9 mIoU improvement on ADE20K.

7 Ablation Studies

To inspect the effectiveness of UniFormer as the backbone, we ablate each key structure design and evaluate the performance on image and video classification datasets. Furthermore, for video backbones, we explore the vital designs of pre-training, training and testing. Finally, we demonstrate the efficiency of our adaption on downstream tasks, and the effectiveness of H-UniFormer.

We conduct ablation studies of the vital components in Table XV.

FFN. As mentioned in Section 3.2, our UniFormer blocks in the shallow layers are instantiated as a transformer-style MobileNet block with extra FFN as in ViT . Hence, we first investigate its effectiveness by replacing our UniFormer blocks in the shallow layers with MobileNet blocks . BN and GELU are added as the original paper, but the expand ratios are set to 3 for similar parameters. Note that the dynamic position embedding is kept for a fair comparison. As expected, our UniFormer outperforms such MobileNet block both in ImageNet (+0.3%) and Kinetics-400 (+0.7%). It shows that, FFN in our model can further mix token context at each position to boost classification accuracy.

DPE. With dynamic position embedding, our UniFormer obviously improves the top-1 accuracy by +0.5% on ImageNet, but +1.7% on Kinetics-400. It shows that via encoding the position information, our DPE can maintain spatial and temporal order, thus contributing to better representation learning, especially for video.

7.2 Pre-training, training and testing for video backbone

In this section, we explore more designs for our video backbones. Firstly, to load 2D pre-trained backbones, it is essential to determine how to inherit self-attention and inflate convolution filters. Hence, we compare the transfer learning performance of different MHRA configurations and inflating methods of filters. Besides, since we use dense sampling for Kinetics, we should confirm the appropriate sampling stride. Furthermore, as we utilize Kinetics pre-trained models for SthSth, it is interesting to explore the effect of sampling methods and dataset scales for pre-trained models. Finally, we ablate the testing strategies for different datasets.

Inflating methods. As indicated in I3D , we can inflate the 2D convolutional filters for easier optimization. Here we consider whether or not to inflate the filters. Note that the first convolutional filter in the patch stem is always inflated for temporal downsampling. As shown in Table XVII, inflating the filters to 3D achieves similar results on Kinetics-400, but obtains performance improvement on SthSth V1. We argue that Kinetics-400 is a scene-related dataset, thus 2D convolution is enough to recognize the action. In contrast, SthSth V1 is a temporal-related dataset, which requires powerful spatiotemporal modeling. Hence, we inflate all the convolutional filters to 3D for better generality by default.

Sampling stride. For dense sampling strategy, the basic hyperparameter is the sampling stride of frames. Intuitively, a larger sampling stride will cover a longer frame range, which is essential for better video understanding. In Table XVIII, we show more results on Kinetics under different sampling strides. As expected, larger sampling stride (i.e. sparser sampling) often achieves higher single-clip results. However, when testing with multi clips, sampling with a frame stride of 4 always performs better.

Sampling methods of Kinetics pre-trained model. For SthSth, we uniformly sample frames as suggested in . Since we load Kinetics pre-trained models for fast convergence, it is necessary to find out whether pre-trained models that cover more frames can help fine-tuning. Table XIX shows that, different pre-trained models achieve similar performances for fine-tuning. We apply 16×\times4 pre-training for better generalization.

Pre-trained dataset scales. In Figure 8, we show more results on SthSth with Kinetics-400/600 pre-training. For UniFormer-S, Kinetics-600 pre-training consistently performs better than Kinetics-400 pre-training, especially for large benchmark SthSth V2. However, both of them achieve comparable results for UniFormer-B. These results indicate that small models are harder to converge and eager for larger dataset pre-training, but big models are not.

Testing strategies. We evaluate our network with various numbers of clips and crops for the validation videos on different datasets. As shown in Figure 8, since Kinetics is a scene-related dataset and trained with dense sampling, multi-clip testing is preferable to cover more frames for boosting performance. Alternatively, Something-Something is a temporal-related dataset and trained with uniform sampling, so multi-crop testing is better for capturing the discriminative motion for boosting performance.

7.3 Adaption designs for downstream tasks

We verify the effectiveness of our adaption for dense prediction tasks in Table XXII, Table XXII and Table XXII. ‘W’, ‘H’ and ‘G’ refer to window, hybrid and global UniFormer style in Stage3 respectively. Note that the pre-trained global UniFormer block can be seen as a window UniFormer block with a large window size, thus the minimal window size in our experiments is 224/32=14.

Table XXII shows results on object detection. Though the hybrid style performs worse than the global style with the 1×\times schedule, it achieves comparable results with the 3×\times schedule, which indicates that training more epochs can narrow the performance gap. We further conduct experiments on semantic segmentation with different model variants in Table XXII. As expected, large window size and global UniFormer blocks contribute to better performances, especially for big models. Moreover, when testing with multi-scale inputs, hybrid style with a window size of 32 obtains similar results to the global style. As for human pose estimation (Table XXII), due to the small input resolution, i.e. 384×\times288, utilizing window style requires more computation for zero paddings. We simply apply global UniFormer blocks for better computation-accuracy balance.

7.4 Designs for H-UniFormer

We further explore the lightweight designs based on UniFormer-XXS160 in Table XXIII. Firstly, we try to remove the score token (i.e., s\mathbf{s}) and simply use the mean of global similarity Anglobal(X,X){\rm A}_{n}^{global}(\mathbf{X},\mathbf{X}) to measure the token importance. Results show that the learnable score token is more helpful for token selection. Besides, the running mean of the similarity score (i.e., A=(A+Apre)/2\mathbf{A}=(\mathbf{A}+\mathbf{A}^{pre})/2) will improve the top-1 accuracy, which verifies the effectiveness of consistent important tokens. Finally, we ablate different shrinking ratios, where we use the ratio of 0.5 for a better trade-off.

8 Visualizations

In Figure 10, we further conduct visualization on validation datasets for various downstream tasks. Such robust qualitative results demonstrate the effectiveness of our UniFormer backbones.

Conclusion

In this paper, we propose a novel UniFormer for efficient visual recognition, which can effectively unify convolution and self-attention in a concise transformer format to overcome redundancy and dependency. We adopt local MHRA in shallow layers to largely reduce computation burden and global MHRA in deep layers to learn global token relation. Extensive experiments demonstrate the powerful modeling capacity of our UniFormer. Via simple yet effective adaption, our UniFormer achieves state-of-the-art results on a broad range of vision tasks with less training cost.

References