TransReID: Transformer-based Object Re-Identification
Shuting He, Hao Luo, Pichao Wang, Fan Wang, Hao Li, Wei Jiang
Introduction
Object re-identification (ReID) aims to associate a particular object across different scenes and camera views, such as in the applications of person ReID and vehicle ReID. Extracting robust and discriminative features is a crucial component of ReID, and has been dominated by CNN-based methods for a long time .
By reviewing CNN-based methods, we find two important issues which are not well addressed in the field of object ReID. (1) Exploiting the rich structural patterns in a global scope is crucial for object ReID . However, CNN-based methods mainly focus on small discriminative regions due to a Gaussian distribution of effective receptive fields . Recently, attention modules have been introduced to explore long-range dependencies , but most of them are embedded in the deep layers and do not solve the principle problem of CNN. Thus, attention-based methods still prefer large continuous areas and are hard to extract multiple diversified discriminative parts (see Figure 1). (2) Fine-grained features with detail information are also important. However, the downsampling operators (e.g. pooling and strided convolution) of CNN reduce spatial resolution of output feature maps, which greatly affect the discrimination ability to distinguish objects with similar appearances . As shown in Figure 2, the details of the backpack are lost in CNN-based feature maps, making it difficult to differentiate the two people.
Recently, Vision Transformer (ViT) and Data-efficient image Transformers (DeiT) have shown that pure transformers can be as effective as CNN-based methods on feature extraction for image recognition. With the introduction of multi-head attention modules and the removal of convolution and downsampling operators, transformer-based models are suitable to solve the aforementioned problems in CNN-based ReID for the following reasons. (1) The multi-head self-attention captures long range dependencies and drives the model to attend diverse human-body parts than CNN models (e.g. thighs, shoulders, waist in Figure 1). (2) Without downsampling operators, transformer can keep more detailed information. For example, one can observe that the difference on feature maps around backpacks (marked by red boxes in Figure 2) can help the model easily differentiate the two people. These advantages motivate us to introduce pure transformers in the object ReID.
Despite its great advantages as discussed above, transformers still need to be designed specifically for object ReID to tackle the unique challenges, such as the large variations (e.g. occlusions, diversity of poses, camera perspective) in images. Substantial efforts have been devoted to alleviating this challenge in CNN-based methods. Among them, local part features and side information (such as cameras and viewpoints) , have been proven to be essential and effective to enhance the feature robustness. Learning part/stripe aggregated features makes it robust against occlusions and misalignments . However, extending the rigid stripe part methods from CNN-based methods to pure transformer-based methods may damage long-range dependencies due to global sequences splitting into several isolated subsequences. In addition, taking side information into consideration, such as camera and viewpoint-specific information, an invariant feature space can be constructed to diminish bias brought by side information variations. However, the complex designs for side information built on CNN, if directly applied to transformers, cannot make full use of the inherent encoding capabilities of transformers. As a result, specific designed modules are inevitable and essential for a pure transformer to successfully handle these challenges.
Therefore, we propose a new object ReID framework dubbed TransReID to learn robust feature representations. Firstly, by making several critical adaptations, we construct a strong baseline framework based on a pure transformer.
Secondly, in order to expand long-range dependencies and enhance feature robustness, we propose a jigsaw patches module (JPM) by rearranging the patch embeddings via shift and shuffle operations and re-grouping them for further feature learning. The JPM is employed on the last layer of the model to extract robust features in parallel with the global branch which does not include this special operation. Hence, the network tends to extract perturbation-invariant and robust features with global context. Thirdly, to further enhance the learning of robust features, a side information embedding (SIE) is introduced. Instead of the special and complex designs in CNN-based methods for utilizing these non-visual clues, we propose a unified framework that effectively incorporates non-visual clues through learnable embeddings to alleviate the data bias brought by cameras or viewpoints. Taking cameras for example, the proposed SIE helps address the vast pairwise similarity discrepancy between inter-camera and intra-camera matching (see Figure 6). SIE can also be easily extended to include any non-visual clues other than the ones we have demonstrated.
To our best knowledge, we are the first to investigate the application of pure transformers in the field of object ReID. The contributions of the paper are summarised:
We propose a strong baseline that exploits the pure transformer for ReID tasks for the first time and achieve comparable performance with CNN-based frameworks.
We design a jigsaw patches module (JPM), consisting of shift and patch shuffle operation, which facilitates perturbation-invariant and robust feature representation of objects.
We introduce a side information embeddings (SIE) that encodes side information by learnable embeddings, and is shown to effectively mitigate the bias of learned features.
The final framework TransReID achieves state-of-the-art performance on both person and vehicle ReID benchmarks including MSMT17, Market-1501, DukeMTMC-reID, Occluded-Duke, VeRi-776 and VehicleID.
Related Work
The studies of object ReID have been mainly focused on person ReID and vehicle ReID, with most state-of-the-art methods based on the CNN structure. A popular pipeline for object ReID is to design suitable loss functions to train a CNN backbone (e.g. ResNet ), which is used to extract features of images. The cross-entropy loss (ID loss) and triplet loss are most widely used in the deep ReID. Luo et al. proposed the BNNeck to better combine ID loss and triplet loss. Sun et al. proposed a unified perspective for ID loss and triplet loss.
Fine-grained Features. Fine-grained features have been learned to aggregate information from different part/region. The fine-grained parts are either automatically generated by roughly horizontal stripes or by semantic parsing. Methods like PCB , MGN , AlignedReID++ , SAN , etc., divide an image into several stripes and extract local features for each stripe. Using parsing or keypoint estimation to align different parts or two objects has also been proven effective for both person and vehicle ReID .
Side Information. For images captured in a cross-camera system, large variations exist in terms of pose, orientation, illumination, resolution, etc. caused by different camera setup and object viewpoints. Some works use side information such as camera ID or viewpoint information to learn invariant features. For example, Camera-based Batch Normalization (CBN) forces the image data from different cameras to be projected onto the same subspace, so that the distribution gap between inter- and intra- camera pairs is largely diminished. Viewpoint/Orientation-invariant feature learning is also important for both person and vehicle ReID.
2 Pure Transformer in Vision
The Transformer model is proposed in to handle sequential data in the field of natural language processing (NLP). Many studies also show its effectiveness for computer-vision tasks. Han et al. and Salman et al. have surveyed the application of the Transformer in the field of computer vision.
Pure Transformer models are becoming more and more popular. For example, Image Processing Transformer (IPT) takes advantage of transformers by using large scale pre-training and achieves the state-of-the-art performance on several image processing tasks like super-resolution, denoising and de-raining. ViT is proposed recently which applies a pure transformer directly to sequences of image patches. However, ViT requires a large-scale dataset to pretrain the model. To overcome this shortcoming, Touvron et al. propose a framework called DeiT which introduces a teacher-student strategy specific for transformers to speed up ViT training without the requirement of large-scale pretraining data.
Methodology
Our object ReID framework is based on transformer-based image classification, but with several critical improvements to capture robust feature (Sec. 3.1). To further boost the robust feature learning in the context of transformer, a jigsaw patch module (JPM) and a side information embeddings (SIE) are carefully devised in Sec. 3.2 and Sec. 3.3. The two modules are jointly trained in an end-to-end manner and shown in Figure 4.
Overlapping Patches. Pure transformer-based models (e.g. ViT, DeiT) split the images into non-overlapping patches, losing local neighboring structures around the patches. Instead, we use a sliding window to generate patches with overlapping pixels. Denoting the step size as , size of the patch as (e.g. ) , then the shape of the area where two adjacent patches overlap is . An input image with a resolution will be split into patches.
where is the floor function and is set smaller than . and represent the numbers of splitting patches in height and width, respectively. The smaller is, the more patches the image will be split into. Intuitively, more patches usually bring better performance with the cost of more computations.
Position Embeddings. As the image resolution for ReID tasks may be different from the original one in image classification, the position embedding pretrained on ImageNet cannot be directly loaded here. Therefore, a bilinear 2D interpolation is introduced to help handle any given input resolution. Similar to ViT, the position embedding is also learnable.
Supervision Learning. We optimize the network by constructing ID loss and triplet loss for global features. The ID loss is the cross-entropy loss without label smoothing. For a triplet set , the triplet loss with soft-margin is shown as follows:
2 Jigsaw Patch Module
Although transformer-based strong baseline can achieve impressive performance in object ReID, it utilizes information from the entire image for object ReID. However, due to challenges like occlusions and misalignments, we may only have partial observation of an object. Learning fine-grained local features such as striped features has been widely used for CNN-based methods to tackle these challenges.
Suppose the hidden features input to the last layer are denoted as . To learn fine-grained local features, a straightforward solution is splitting into groups in order which concatenate the shared token and then feed feature groups into a shared transformer layer to learn local features denoted as and is the output token of -th group. But it may not take full advantage of global dependencies for the transformer because each local segment only considers a part of the continuous patch embeddings.
To address the aforementioned issues, we propose a jigsaw patch module (JPM) to shuffle the patch embeddings and then re-group them into different parts, each of which contains several random patch embeddings of an entire image. In addition, extra perturbation introduced in training also helps improve the robustness of object ReID model. Inspired by ShuffleNet , the patch embeddings are shuffled via a shift operation and a patch shuffle operation. The sequences embeddings are shuffled as follow:
Step1: The shift operation. The first patches (except for [cls] token) are moved to the end, i.e. is shifted in steps to become .
Step2: The patch shuffle operation. The shifted patches are further shuffled by the patch shuffle operation with groups. The hidden features become .
With the shift and shuffle operation, the local feature can cover patches from different body or vehicle parts which means that the local features hold global discriminative capability.
As shown in Figure 4, paralleling with the jigsaw patch, another global branch which is a standard transformer encodes into , where is served as the global feature of CNN-based methods. Finally, the global feature and local features are trained with and . The overall loss is computed as follow:
During inference, we concatenate the global feature and local features as the final feature representation. Using only is a variation with lower computational cost and slight performance degradation.
3 Side Information Embeddings
After obtaining fine-grained feature representations, features are still susceptible to camera or viewpoint variations. In other words, the trained model may easily fail to distinguish the same object from different perspectives due to scene-bias. Therefore, we propose a Side Information Embedding (SIE) to incorporate the non-visual information, such as cameras or viewpoints, into embedding representations to learn invariant features.
Finally, the input sequences with camera ID and viewpoint ID are fed into transformer layers as follows:
where is the raw input sequences in Eq. 2 and is a hyperparameter to balance the weight of SIE. As the position embeddings are different for each patch but the same across different images, and are the same for each patch but may have different values for different images. Transformer layers are able to encode embeddings with different distribution properties which can then be added directly.
Here we have only demonstrate the usage of SIE with camera and viewpoint information which are both categorical variables. In practice, SIE can be further extended to encode more kinds of information, including both categorical and numerical variables. In our experiments on different benchmarks, camera and viewpoint information is included wherever available.
Experiments
We evaluate our proposed method on four person ReID datasets, Market-1501 , DukeMTMC-reID , MSMT17 , Occluded-Duke , and two vehicle ReID datasets, VeRi-776 and VehicleID . It is noted that, unlike other datasets, images in Occluded-Duke are selected from DukeMTMC-reID and the training/query/gallery set contains 9%/ 100%/ 10% occluded images respectively. All datasets except VehicleID provide camera ID for each image, while only VeRi-776 and VehicleID dataset provide viewpoint labels for each image. The details of these datasets are summarized in Table 1.
2 Implementation
Unless otherwise specified, all person images are resized to and all vehicle images are resized to . The training images are augmented with random horizontal flipping, padding, random cropping and random erasing . The batch size is set to 64 with 4 images per ID. SGD optimizer is employed with a momentum of 0.9 and the weight decay of 1e-4. The learning rate is initialized as 0.008 with cosine learning rate decay. Unless otherwise specified, we set and for person and vehicle ReID datasets, respectively.
All the experiments are performed with one Nvidia Tesla V100 GPU using the PyTorch toolbox http://pytorch.org with FP16 training . The initial weights of ViT are pre-trained on ImageNet-21K and then finetuned on ImageNet-1K, while the initial weights of DeiT are trained only on ImageNet-1K.
Evaluation Protocols. Following conventions in the ReID community, we evaluate all methods with Cumulative Matching Characteristic (CMC) curves and the mean Average Precision (mAP).
3 Results of Transform-based Baseline
In this section, we compare CNN-based and transformer-based backbones in Table 2. To show the trade-off between computation and performance, several different backbones are chosen. DeiT-small, DeiT-Base, ViT-Base denoted as DeiT-S, DeiT-B, ViT-B, respectively. ViT-B/16s=14 means ViT-Base with patch size 16 and step size in overlapping patches setting. For a comprehensive comparison, inference time consumption of each backbone is included as well.
We can observe a large gap in model capacity between the ResNet series and DeiT/ViT. DeiT-S/16 is a little bit better in performance and speed compared to ResNet50. DeiT-B/16 and ViT-B/16 achieve similar performance with ResNeSt50 backbone, with less inference time than ResNeSt50 (1.79x vs 1.86x). When we reduce the step size of the sliding window , the performance of the Baseline can be improved while the inference time is also increasing. ViT-B/16s=12 is faster than ResNeSt200 (2.81x vs 3.12x) and performs slightly better than ResNeSt200 on ReID benchmarks. Therefore, ViT-B/16s=12 achieves better speed-accuracy trade-off than ResNeSt200. In addition, we believe that DeiT/ViT still have lots of room for improvement in terms of computational efficiency.
4 Ablation Study of JPM
The effectiveness of the proposed JPM module is validated in Table 3. JPM provides +2.6% mAP and +1.0% mAP improvements compared to baseline on MSMT17 and VeRi-776, respectively. Increasing the number of groups can improve the performance while slightly increasing inference time. In our experiment, is a choice to trade off speed and performance. Comparing JPM and JPM w/o rearrange, we can observe that the shift and shuffle operation helps the model learn more discriminative features with +0.5% mAP and +0.2% mAP improvements on MSMT17 and VeRi-776, respectively. It is also observed that, if only the global feature is used in inference stage (still trained with full JPM), the performance (denoted as “w/o local”) is nearly comparable with the version of full set of features, which suggests us to only use the global feature as an efficient variation with lower storage cost and computational cost in the inference stage. The attention maps visualized in Figure 5 show that JPM with the rearrange operation can help the model learn more global context information and more discriminative parts, which makes the model more robust to perturbations.
5 Ablation Study of SIE
Performance Analysis. In Table 4, we evaluate the effectiveness of the SIE on MSMT17 and VeRi-776. MSMT17 does not provide viewpoint annotations, so the results of SIE which only encode camera information are shown for MSMT17. VeRi-776 not only have a camera ID of each image, but is also annotated with 8 different viewpoints according to vehicle orientation. Therefore, the results are shown with SIE encoding various combinations of camera ID and/or viewpoints information.
When SIE encodes only the camera IDs of images, the model gains 1.4% mAP and 0.1% rank-1 accuracy improvements on MSMT17. Similar conclusion can be made on VeRi-776. Baseline obtains 78.5% mAP when SIE encodes viewpoint information. The accuracy increases to 79.6% mAP when both camera IDs and viewpoint labels are encoded at the same time. If the encoding is changed to , which is sub-optimal as discussed in Section 3.3, we can only achieve 78.3% mAP on VeRi-776. Therefore, the proposed is a better encoding manner.
Visualization of Distance Distribution. As shown in Figure 6, the distribution gaps with cameras and viewpoints variations are obvious in Figure 6(a) and Figure 6(b), respectively. When we introduce the SIE module into Baseline, the distribution gaps between inter-camera/viewpoint and intra-camera/viewpoint are reduced, which shows that the SIE module weakens the negative effect of the scene-bias caused by various cameras and viewpoints.
Ablation Study of . We analyze the influence of weight of the SIE module on the performance in Figure 7. When , Baseline achieves 61.0% mAP and 78.2% mAP on MSMT17 and VeRi-776, respectively. With increasing, the mAP is improved to 63.0% mAP ( for MSMT17) and 79.9% mAP ( for VeRi-776), which means the SIE module now is beneficial for learning invariant features. Continuing to increase , the performance is degraded because the weights for feature embedding and the position embedding are weakened.
6 Ablation Study of TransReID
Finally, we evaluate the benefits of introducing JPM and SIE in Table 5. For the Baseline, JPM and SIE improve the performance by +2.6%/+1.0% mAP and +1.4%/+1.4% mAP on MSMT17/VeRi-776, respectively. With these two modules used together, TransReID achieves 64.9% (+3.9%) mAP and 80.6% (+2.4%) mAP on MSMT17 and VeRi-776, respectively. The experimental results show the effectiveness of our proposed JPM, SIE, and the overall framework.
7 Comparison with State-of-the-Art Methods
In Table 6, our TransReID is compared with state-of-the-art methods on six benchmarks including person ReID, occluded ReID and vehicle ReID.
Person ReID. On MSMT17 and DukeMTMC-reID, TransReID∗ (DeiT-B/16) outperforms the previous state-of-the-art methods by a large margin (+5.5%/+2.1% mAP). On Market-1501, TransReID∗ (256128) achieves comparable performance with state-of-the-art methods especially on mAP. Our method also shows superiority when compared with methods which also integrate camera information like CBN .
Occluded ReID. ISP implicitly uses human body semantic information through iterative clustering and HOReID introduces external pose models to align body parts. TransReID (DeiT-B/16) achieves 55.6% mAP with a large margin improvement (at least +3.3% mAP) compared to aforementioned methods, without requiring any semantic and pose information to align body parts, which shows the ability of TransReID to generate robust feature representations. Furthermore, TransReID∗ improves the performance to 58.1% mAP with the help of overlapping patches.
Vehicle ReID. On VeRi-776, TransReID∗ (DeiT-B/16) reaches 82.3% mAP surpassing GLAMOR by 2.0% mAP. When only using viewpoint annotations, TransReID∗ still outperforms VANet and SAVER on both VeRi-776 and VehicleID. Our method achieves state-of-the-art performance about 85.2% Rank-1 accuracy on VehicleID.
DeiT vs ViT vs CNN. TransReID∗ (DeiT-B/16) reaches competitive performance with existing methods under a fair comparison (ImageNet-1K pre-training). Extra results of our methods with ViT-B/16 are also reported in Table 6 for further comparison. DeiT-B/16 achieves similar performance with ViT-B/16 for shorter image patch sequences. When the number of input patches is increasing, ViT-B/16 reaches better performance than DeiT-B/16, which shows ImageNet-21K pre-training provides ViT-B/16 better generalization capability. Although CNN-based methods mainly report performance with the ResNet50 backbone, they may include multiple branches, attention modules, semantic models, or other modules that increase computational consumption. We have conducted a fair comparison on inference speed between TransReID∗ and MGN on the same computing hardware. Compared with MGN, TransReID* is 4.8% faster in speed. Therefore, TransReID* can achieve more promising performance under comparable computation to most of CNN-based methods.
Conclusion
In this paper, we investigate a pure transformer framework for the object ReID task, and propose two novel modules, i.e., jigsaw patch module (JPM) and side information embedding (SIE). The final framework TransReID outperforms all other state-of-the-art methods by a large margin on several popular person/vehicle ReID datasets including MSMT17, Market-1501, DukeMTMC-reID, Occluded-Duke, VeRi-776 and VehicleID. Based on the promising results achieved by TransReID, we believe the transformer has great potential to be further explored for ReID tasks. Based on the rich experience gained from CNN-based methods, it is in prospect that more efficient transformer-based networks can be designed with better representation power and less computational cost.
References
Appendix
Appendix A More Experimental Results
A transformer-based strong baseline with a few critical improvements has been introduced in Section 3.1 of the main paper. In this section, hyper-parameters and the settings for training such a baseline model will be analyzed in detail. Ablation studies are shown in Table 7 for performance on MSMT17 and Veri-776 with different variations of the training settings.
Initialization and hyper-parameters. For our experiments, we initialize the pure transformer with ViT or DeiT ImageNet pre-trained weights and we initialize the weights for the SIE with a truncated normal distribution . Compared with ViT, DeiT is more sensitive to hyper-parameter settings. For the training of DeiT, we use a learning rate of 0.05 on MSMT17 and a high random erasing probability with 0.8 on each dataset to avoid overfitting. Other hyper-parameters settings are the same with ViT.
Optimizer. Transformers are sensitive to the choice of the optimizer. Directly applying Adam optimizer with the hyper-parameters commonly used in ReID community to transformer-based models will cause a significant drop in performance. AdamW is a commonly used optimizer for training transformer-based models, with much better performance compared with Adam. The best results are actually achieved by SGD in our experiments.
Network Configuration. Position embeddings incorporate crucial spatial information which provides a significant boost in performance and is one of the key ingredients of our proposed training procedure. Without the position embeddings, the performance decreases by 38.6% mAP and 10.2% mAP on MSMT17 and VeRi-776, respectively.
Introducing stochastic depth can boost the mAP performance by about 1%, and it has also been proved to facilitate the convergence of transformer, especially for those deep ones . Regarding other regularization methods, adding either drop out or attention drop out will result in performance drop. In our experiments, we set all the probability of regularization methods as 0.1.
Loss Function. Different choices of loss functions have been compared in the bottom section of Table 7. The soft version of triplet loss provides 0.7% mAP improvement on MSMT17 compared with the regular triplet loss. Introducing label smoothing is harmful to performance, even though it has been a widely adopted trick. Therefore, the best combination for loss functions is soft triplet loss and cross entropy loss without label smoothing.
A.2 More Ablation Studies of JPM and SIE
In the main paper, we have demonstrated the effectiveness of using JPM and SIE based on the Baseline (ViT-B/16). More results about JPM and SIE are shown in Table 8 and Table 9 respectively, with the Baseline ViT-B/16s=12, which is supposed to have better feature representation ability and higher performance than ViT-B/16. From Table 8, we observe that: (1) The proposed JPM performs better with the rearrange schemes, indicating that the shift and patch shuffle operation help the model learn more discriminative features which are robust against perturbations. (2) The JPM module provides a consistent performance improvement over the baselines, no matter the baseline is ViT-B/16 or the stronger ViT-B/16s=12, demonstrating the effectiveness of the proposed JPM.
Similar conclusions can be made from Table 9. (1) We make better use of the viewpoint and camera information so that they are complementary with each other and combining them leads to the best performance. (2) Introducing SIE provides consistent improvement over the baselines of either ViT-B/16 or ViT-B/16s=12.
Appendix B Analysis on Rearranging Patches in JPM
Although transformers can capture the global information in the image very well, a patch token still has a strong correlation with the corresponding patch. ViT-FRCNN shows that the output embeddings of the last layer can be reshaped as a spatial feature map that includes location information. In other words, if we directly divide the original patch embeddings into parts, each part may only consider a part of the continuous patch embeddings. Therefore, to better capture the long-range dependencies, we rearrange the patch embeddings and then re-group them into different parts, each of which contains several random patch embeddings of an entire image. In this way, the JPM module help to learn robust features with improved discrimination ability and more diversified coverage.
To verify the above point, we visualize the learned attention of local features ( in our cases) by JPM module in Figure 8. Brighter region means higher corresponding weights. Several observations can be made from Figure 8: (1) The attention learned by the “JPM w/o rearrange” tends to focus on limited receptive fields (i.e. the range of the corresponding patch sequences) due to global sequences being split into several isolated sub-sequences. For example, “Part 1” mainly pays attention to the head of a person, and “Part 4” is mainly focused around the bottom area. (2) In contrast, “JPM w/ rearrange” is able to capture long-range dependencies and each part has attention responses across the whole image because it is forced to extend its scope to the whole image through the rearranging operation. (3) According to the superior ReID performance and the intuitive visualization of rearranging effect, JPM is proved to not only capture more details at finer granularities but also learn robust and discriminative representations in the global context.
Appendix C More Visualization of Attention Maps
In the main paper, we use Grad-CAM to visualize the gradient responses of our schemes, CNN-based methods, and CNN+attention methods. Following the similar setup, Figure 9 shows more visualization results, with the similar conclusion that transformer-based methods capture global context information and more discriminative parts, which are further enhanced in our proposed TransReID for better performance.