S$^2$-MLPv2: Improved Spatial-Shift MLP Architecture for Vision
Tan Yu, Xu Li, Yunfeng Cai, Mingming Sun, Ping Li
Introduction
Recently, extensive studies on computer vision are conducted to achieve high performance with less inductive bias. Two types of architectures emerge including vision Transformers (Dosovitskiy et al., 2021; Touvron et al., 2020) and MLP-based backbones (Tolstikhin et al., 2021; Touvron et al., 2021a). Compared with de facto vision backbone CNN (He et al., 2016) with delicately devised convolution kernels, both vision Transformers and MLP-based backbones have achieved competitive performance in image recognition without expensive hand-crafted design. Specifically, vision Transformer models stack a series of Transformer blocks, achieving the global reception field.
MLP-based methods such as MLP-Mixer (Tolstikhin et al., 2021) and ResMLP (Touvron et al., 2021a) achieve the communication between patches through projections along different patches implemented by MLP. Different from MLP-Mixer and ResMLP, spatial-shift MLP (S2-MLP) (Yu et al., 2021b) adopts a very straightforward operation, spatial shifting, for communications between patches, achieving higher image recognition accuracy on ImageNet1K dataset without external training data. In parallel, Vision Permutator (ViP) (Hou et al., 2021) encodes the feature representation along the height and width dimensions and meanwhile exploits the finer patch size with a two-level pyramid structure, achieving better performance than S2-MLP. CCS-MLP (Yu et al., 2021a) devises a circulant token-mixing MLP for achieving the translation-invariance property. Global Filter Networks (GFNet) (Rao et al., 2021b) exploits 2D Fourier Transform to map the spatial patch features into the frequency domain and conducts the cross-patch communications in the frequency domain. As pointed by Rao et al. (2021b), the token-mixing operation in the frequency domain is equivalent to depthwise convolution with circulant weights. To achieve a high recognition accuracy, GFNet also utilizes patches of smaller size with a pyramid structure. More recently, AS-MLP (Lian et al., 2021) axially shifts channels of the feature map and devises a four-level pyramid, achieving excellent performance. In parallel, Cycle-MLP (Chen et al., 2021a) devises several pseudo-kernels for spatial projection and also achieves outstanding performance. It is worth noting that, both AS-MLP (Lian et al., 2021) and Cycle-MLP (Chen et al., 2021a) are based on the well-devised four-level pyramid.
In this work, we rethink the design of spatial-shift MLP (S2-MLP) (Yu et al., 2021b) and propose an improved spatial-shift MLP (S2-MLPv2). Compared with the original S2-MLP, the modifications are mainly conducted on two aspects:
As visualized in Figure 1 (b), we expand the feature map along the channel dimension and split the expanded feature map into multiple parts. For different parts, we conduct different spatial-shift operations. We exploit the split-attention operation (Zhang et al., 2020) to fuse these split parts.
We adopt smaller-scale patches and the hierarchical pyramid structure like existing MLP-based architectures such as ViP (Hou et al., 2021), GFNet (Rao et al., 2021b), AS-MLP (Lian et al., 2021) and Cycle-MLP (Chen et al., 2021a).
We term the improved spatial-shift MLP architecture as S2-MLPv2. We visualize the difference between the original spatial-shift MLP (S2-MLP) and the improved verision, S2-MLPv2, in Figure 1. Our experiments conduct on the public benchmark, ImageNet-1K, demonstrates the state-of-the-art image recognition accuracy of the proposed S2-MLPv2. Specifically, using 55M parameters, our medium-scale model, S2-MLPv2-Medium achieves top-1 accuracy using images without self-attention and external training data.
Related Work
vision Transformer. vision Transformer (ViT) (Dosovitskiy et al., 2021) crops an image into patches, and treat each patch as a token in the input of Transformer. These patches/tokens are processed by a stack of Transformer layers for communications with each other. It has achieved competitive image recognition accuracy as CNNs using huge-scale pre-training datasets. DeiT (Touvron et al., 2020) adopts more advanced optimizer as well as data augmentation methods, achieving promising results using medium-scale pre-training datasets. Pyramid vision transformer (PvT) (Wang et al., 2021b) and PiT (Heo et al., 2021) exploit a pyramid structure which gradually shrinks the spatial dimension and expands the hidden size, achieving better performance. Tokens-to-Token (T2T) (Yuan et al., 2021) and Transformer-in-Transformer (TNT) (Han et al., 2021) improve the effectiveness of modeling the local structure of each patch/token. To overcome the inefficiency of the global self-attention, Swin (Liu et al., 2021b) conducts the self-attention within local windows but achieves the global reception field through shifting the window settings. Shuffle Transformer (Huang et al., 2021) also exploits the local self-attention windows and achieves the cross-window communications through switching the spatial dimension and the feature dimension. Twins (Chu et al., 2021a) enhances the self-attention within local windows by the global sub-sampled attention. DynamicViT (Rao et al., 2021a) and SViTE (Chen et al., 2021b) exploit the sparsity for achieving high efficiency. CaiT (Touvron et al., 2021b) explores the extremely deep architecture by stacking tens of layers. Recently, PVTv2 (Wang et al., 2021a) improves PvT using overlapping patch embedding, convolutional feedforward networks, and linear-complexity attention layers. CSwin Transformer (Dong et al., 2021) improves Swin through cross-shaped windows computing self-attention in the horizontal and vertical stripes in parallel. Focal Transformer (Yang et al., 2021) also develops more advanced local windows which attend fine-grain tokens only locally, but the summarized ones globally.
MLP-based architectures. MLP-mixer (Tolstikhin et al., 2021) is the pioneering work for MLP-based vision backbone. It proposes a token-mixing MLP consisting of two fully-connected layers for communications between patches. Res-MLP (Touvron et al., 2021a) simplifies the token-mixing MLP to a single fully-connected layer and explores the deeper architecture with more layers. Spatial-shift MLP backbone (S2-MLP) (Yu et al., 2021b) adopts the spatial-shift operation for cross-patch communications. Vision Permutator (ViP) (Hou et al., 2021) mixes tokens along the height dimension and the width dimension, separately. Meanwhile, ViP adopts a pyramid structure as PvT (Wang et al., 2021b) and achieves considerably better performance than MLP-mixer, Res-MLP and S2-MLP. CCS-MLP (Yu et al., 2021a) rethinks the design of token-mixing MLP in MLP-mixer and Res-MLP, and propose a circulant channel-specific MLP. Specifically, they devise the weight matrix of token-mixing MLP as a circulant matrix, taking fewer parameters. Meanwhile, the multiplication between vector and circulant matrix can be efficiently computed through Fast Fourier Transform (FFT). Global Filter Network (GFNet) (Rao et al., 2021b) maps the patch features to the frequency domain through 2D FFT and mixes the tokens in the frequency domain. As proved by Rao et al. (2021b), the global filter in GFNet is equivalent to a depthwise global circular convolution with the filter size H W. Meanwhile, GFNet also exploits pyramid structure for boosting the recognition accuracy. In this work, we rethink the design of S2-MLP and considerably improves its performance in image recognition.
Preliminary
In this section, we briefly review the structure of S2-MLP (Yu et al., 2021b) architecture. It consists of the patch embedding layer, a stack of S2-MLP blocks and the classification head.
Patch embedding layer. It first crops an image of size into patches. Each patch is of size and . It then maps each patch into a -dimensional vector through a fully-connected layer.
It is worth noting that, S2-MLP (Yu et al., 2021b) stacks Spatial-shift MLP blocks with the same settings and does not exploit pyramid structure as its MLP-backone counterparts such as Vision Permutator (Hou et al., 2021) and Global Filter Network (GFNet) (Rao et al., 2021b).
2 Split Attention
Vision Permutator (Hou et al., 2021) adopts split attention proposed in ResNeSt (Zhang et al., 2020) for enhancing multiple feature maps from different operations. Specifically, we denote features maps of the same size by where is the number of patches and is the number of channels, the split-attention operation first averages them and obtains
where denotes the element-wise multiplication between two vectors.
S2-MLPv2
In this section, we introduce the proposed S2-MLPv2 architecture. Similar to S2-MLP backbone, S2-MLPv2 backbone consists of the patch embedding layer, a stack of S2-MLPv2 blocks and the classification head. Since we have introduced the patch embedding layer in the previous section, we only introduce the proposed S2-MLPv2 block below.
The channel-mixing MLP (CM-MLP) adopts the same structure as MLP-mixer (Tolstikhin et al., 2021) and ResMLP (Touvron et al., 2021a) and thus we skip their details here. Below we only introduce the proposed S2-MLPv2 component in detail.
Then it equally splits the expanded feature map along the channel dimension into three parts:
Then the attended feature map is further fed into another MLP layer for generating the output
The structure of the proposed S2-MLP module is visualized in Figure 3 and the details are listed in Algorithm 1.
2 Pyramid structure
Following vision Permutator (Hou et al., 2021), we also exploit the two-level pyramid structure to enhance the performance. To make a fair comparison with Vision Permutator (Hou et al., 2021), we adopt the exact same pyramid structure. The details are in Table 1. We notice that counterpart works such as PVTv2 (Wang et al., 2021a), AS-MLP (Lian et al., 2021), and Cycle-MLP (Chen et al., 2021a) adopt more advanced pyramid structure with smaller patches in the early blocks. The smaller patches might be better at capturing the fine-grained visual details and lead to higher recognition accuracy. Nevertheless, due to the limited computing resources, it is unfeasible for us to re-implement all these pyramid structures. Moreover, we also notice that Vision Permutator devises a large model with considerably more parameters and FLOPs. Nevertheless, due to the limited computing resources, the large model is not feasible for us, either.
Experiments
We testify the proposed S2-MLPv2 on ImageNet-1K dataset (Deng et al., 2009). We do not use external data for training. The implementation is based on the PaddlePaddle deep learning platform.
Implementation details. Following DeiT (Touvron et al., 2020), we adopt AdamW (Loshchilov & Hutter, 2019) as optimizer. We train both the small model and the medium model using four NVIDIA A100 GPU cards. For the small model, we set the batch size as 1024. In contrast, for the medium model, we only set the batch size as 744 due to the GPU memory limitation of four NVIDIA A100 GPU cards. We set the initial learning rate as 2e-3 and decay it to 1e-5 in 300 epochs using a cosine function. The weight decay rate is set to be 5e-2 following previous works (Touvron et al., 2020; Hou et al., 2021). We also conduct warming up in the first 10 epochs following Hou et al. (2021). Moreover, as adopted by Touvron et al. (2020); Hou et al. (2021), we conduct multiple data augmentation methods including Rand-Augment (Cubuk et al., 2020), Mixup (Zhang et al., 2018) and CutMix (Yun et al., 2019) and CutOut (Zhong et al., 2020). Like DeiT (Touvron et al., 2020) and Vision Permutator (Hou et al., 2021), we adopt exponential moving average (EMA) model (Laine & Aila, 2016). Meanwhile, we also use label smoothing (Szegedy et al., 2016) with a smooth ratio of and DropPath (Huang et al., 2016) with a drop ratio of for both small and medium settings.
Influence of each split. As equation 9, we fuse three splits through the split attention. In this section, we evaluate the influence of removing one of them. The experiments are conducted in Small/7 settings. As shown in Table 7, when using only and , the top-1 accuracy drops from to . Meanwhile, when removing , the top-1 accuracy decreases to .
In this paper, we improve the spatial-shift MLP (S2-MLP) and propose an S2-MLPv2 model. It expands the feature map and splits the expanded feature map into three splits. It shifts each split individually and then fuses the split feature maps through split-attention. Meanwhile, we exploit the hierarchical pyramid to improve its capability of modeling fine-grained details for higher recognition accuracy. Using M parameters, our S2-MLPv2-Medium model achieves top-1 accuracy on ImageNet1K dataset using images without external training datasets, which is the state-of-the-art performance among MLP-based methods. Meanwhile, compared with Transformer-based methods, our S2-MLPv2 model has achieved comparable accuracy without self-attention and fewer parameters.
Compared with the pioneering MLP-based works such as MLP-mixer, ResMLP as well as recent MLP-like models including Vision Permutator and GFNet, another important advantage of the spatial-shift MLP is that the shapes of spatial-shift MLPs are invariant to the input scale of images. Thus, the spatial-shift MLP model pre-trained by images of a specific scale can be well adopted for down-stream tasks with various-sized input images.
The future work will be devoted to continuously improving the image recognition accuracy of the spatial-shift MLP architecture. A promising and straightforward direction is to attempt smaller-size patches and the more advanced four-level pyramid as CycleMLP and AS-MLP for further reducing the FLOPs and shortening the recognition gap between the Transformer-based models.