Interaction-and-Aggregation Network for Person Re-identification
Ruibing Hou, Bingpeng Ma, Hong Chang, Xinqian Gu, Shiguang Shan, Xilin Chen
Introduction
Person re-identification (reID) aims at identifying a person of interest across different cameras with a given probe. It plays a significant role in intelligent surveillance systems. In recent years, Deep Convolutional Neural Networks (CNNs), which typically stack convolution and pooling layers to learn discriminative features, have obtained state-of-the-art results for person reID. Despite of years of efforts, there still exist many challenges such as large variations in person pose, scale, and background clutter.
Body part misalignment is a critical influencing factor on reID results, which can be attributed to two causes. First, pedestrians naturally take on various poses as shown in Fig. 1 (a). Second, the body parts have various scales across different images of the same person caused by imperfect pedestrian detection, as illustrated in Fig. 1 (b). To resolve these problems, some approaches have been proposed recently. One way is to localize body parts explicitly and combine the representations over them . This scheme requires highly-accurate part detection. Unfortunately, even state-of-the-art part detection solutions are not perfect. Another type of methods resorts to multi-scale features fusion where the feature maps are computed at multiple layers of a network . Nevertheless, these methods only employ manually specified scales, which are ineffective to model large scale variations. In short, existing methods, which attempt to utilize body part detection or multi-scale features, are still limited in modeling the large variations in body pose and scale.
An essential reason why these approaches are not robust to body pose and scale variations is that they all use CNNs to extract pedestrian features. Actually, CNNs are inherently limited in modeling large geometric transformations. The limitation originates from the fixed geometric structures of CNNs modules: a convolution unit which samples the input feature map at fixed locations and a pooling layer which reduces the spatial resolution at a fixed ratio. There lacks internal mechanisms to handle the body pose and scale variations. For one thing, the receptive fields of the feature maps are pre-defined rectangles, which can not adaptively localize the non-rigid body parts with different poses. For another, the receptive fields of all activation units in the same CNN layer have the same size, which is undesirable for high level CNN layers to encode semantics for body parts of different scales.
In this paper, we propose a new network structure, Interaction-and-Aggregation (IA), to enhance the feature representation capability of CNNs, especially at the presence of body pose and scale variations. IA consists of two modules: Spatial Interaction-and-Aggregation (SIA) and Channel Interaction-and-Aggregation (CIA). Unlike CNNs which extract features with fixed geometric structure, SIA adaptively determines the receptive fields according to the pose and scale of input person image. More specifically, given the intermediate feature maps from CNNs, SIA generates spatial semantic relation maps to discover two types of interdependencies between different image positions: appearance relations where positions with similar feature representations have a higher correlation, and location relations where positions close to each other tend to have a higher correlation. In this way, the body parts with various poses and scales can be adaptively localized. Based on the spatial relation maps, an aggregation operation is adopted to update the feature maps via aggregating the semantically correlated features across different positions. Similar with SIA in principle, we propose CIA to further enhance the representation power of CNNs. Unlike CNNs where the features from different channels are assumed independently, CIA explicitly models the semantic interdependencies between channels. Specially, for small-scale visual cues (e.g. bags) that easily fade away in the high-level features from CNNs, CIA can selectively aggregate the semantically similar features of the visual cues across all channels to manifest their feature representations.
Both modules are computationally lightweight and impose only a slight increase in model complexity. They can be readily inserted into deep CNNs at any depth. In our work, we add IA blocks to ResNet-50 to generate Interaction-and-Aggregation network (IANet) for person reID. We demonstrate the effectiveness of IANet on three reID datasets, and our method outperforms state-of-the-art methods under multiple evaluation metrics.
Related Work
Person re-identification. Person reID methods focus on two key points: learning a powerful feature representation for images and designing an effective distance metric . Recently, deep learning approaches have obtained state-of-art results for reID. We focus our discussion on those which attempt to address the problem of body pose and scale variations.
Body part detection results have been exploited for reID to extract features robust to pose and scale variations. Most approaches attempt to localize body parts explicitly and combine the representations over global features. Specifically, Zhao et al. used a region proposal network, which is trained on an auxiliary pose dataset, to detect body parts. Su et al. proposed a sub-network to estimate the human pose that is used to crop the body parts. Besides, human parsing method and body part specific attention modeling had also been adopted to explicitly alleviate the pose variations problem. However, part detection in low resolution pedestrian images has its own challenges, and the inevitable detection errors could propagate to the subsequent reID task.
Another line of approaches attempts to utilize multi-scale features. Liu et al. and Chen et al. proposed an architecture consisting of multiple branches for learning multi-scale features and one branch for feature fusion. Chen et al. and Shen et al. used the hourglass-like network to generate multi-scale features. Wang et al. and Chang et al. directly fused the feature maps across multiple layers to generate a single feature. Nevertheless, these methods employ pre-defined scales that are limited in modeling large scale variations.
In contrast to the above works that rely on part detection or pre-defined scales, our proposed SIA can adaptively localize the body parts under various poses and scales and aggregate semantic features therein. Therefore, SIA can be easily inserted into existing networks, enhancing their feature representation power.
Modeling geometric variations. There are some works which enhance the feature representation power with respect to geometric variations. Traditional methods include scale invariant feature transform (SIFT) and ORB . A lot of recent works are aimed at CNNs. Some works learn invariant CNN representations with respect to specific transformations such as symmetry , scale and rotation . However, these works assume the transformations are fixed and known, which restricts their generalization to new tasks with unknown transformations. Other works adaptively learn the spatial transformations from data. Spatial Transform Network warped the feature map via a global parametric transformation. The works augmented the sampling locations in the convolution with offsets and learn the offsets via back-propagation end-to-end.
Our work is fundamentally different from those works in two folds. First, the basic idea and formulation are different. The above works usually learn a parametric transformation with large amount of training data, which is infeasible for reID task with a small dataset. Differently, our proposed SIA computes spatial semantic similarities to adaptively aggregate features from same body parts without any parameters. Second, all above works do not take the channel relations into consideration. In contrast, our proposed CIA explicitly models the correlations between channels, which significantly enhances the feature representation power.
Interaction-and-Aggregation Network
In this section, we first introduce SIA and CIA modules, respectively. Then, IA block, which integrates SIA and CIA modules, is illustrated, followed by IANet for person reID. Finally, we provide some discussions on the relationships between the proposed modules and other related models.
With fixed local receptive fields, CNNs are limited in representing person images with large variations in body pose and scale. To address this problem, we design the SIA module to model spatial features interdependencies. SIA could adaptively determine the receptive field for each spatial feature, thus improving the feature robustness to body pose and scale variations.
Appearance Relations. We measure the appearance similarity between any two positions of an input feature map to generate the appearance relation map. Du et al. have pointed out that local features at neighboring spatial positions have high correlation since their receptive fields are often overlapped. So the patches involving neighboring positions could capture more precise appearance. Inspired by their views, we propose to incorporate contextual information for any position in order to obtain more precise appearance similarities.
where and denote the features in the spatial position of patches and , respectively. Notably, the softmax dramatically suppresses small similarity values corresponding to different body parts. Through incorporating context and suppressing dissimilarities, the relation map can roughly localize the body parts under various poses and scales. We call the process single-context interaction as only one patch size is considered, and is the single-context appearance relation map.
As shown in Fig. 4, the relation maps with small context patches (e.g., ) capture more positive regions, but introduce some outliers, e.g., the located regions of the foot contain some positions corresponding to the trunk. The relation maps with large context patches (e.g., ) filter out the outliers, but ignore some positive regions. Therefore, we introduce multi-context interaction by fusing multiple single-context relation maps with different context patch sizes. The multi-context appearance relation map is computed as:
where denotes the number of context levels and is a fusion function with element-wise product. From Fig. 4 (c), multi-context interaction can alleviate both problems and localize the body parts more precisely.
Location Relations. As for pedestrian images, local features corresponding to the same body part are spatially close. To take advantage of the spatial structure information, we introduce location relations, in which features from nearby locations have a higher correlation.
Formally, the location relation between spatial features and is computed via a two-dimensional Gaussian function as follows:
where and denote the location coordinates of features and respectively, and are the standard deviations used to tune the Gaussian function. We then normalize ’s so that the sum of the location relation values connected to equals to . The resulting spatial location relation map is:
We can see that the location relation between and exponentially decreases with the increase of their spatial distance. Notably, is computed based on the spatial structure of the input image, which can constrain and complement the appearance relations.
The spatial semantic relations () integrates the appearance with location relations, which is formulated as:
2 CIA Module
Current reID models typically stack multiple convolution layers to extract pedestrian features. With increasing the number of layers, these models could easily lose small scale visual cues, such as bags and shoes. However, these fine-grained cues are very useful to distinguish the pedestrian pairs with small inter-class variations. Zhang et al. have discovered that most channel maps of high-level features show strong responses for specific parts. Motivated by their views, we build the CIA module to aggregate semantically similar features across all channels, which could enhance the feature representation of specific parts.
3 IA Block
We turn the SIA (CIA) module into SIA (CIA) block that can be easily incorporated into existing architectures. As shown in Fig. 6 (a), SIA (CIA) block is defined as:
where is the input feature map, is the output of SIA or CIA modules that is given in Eq. 6 or Eq. 8, and BN is a batch normalization layer which adjusts the scale of with respect to the input. The residual connection () allows us to insert a new block into any pre-trained model, without breaking its initial performance (e.g. the parameters of BN are initialized to zeros).
Given an input feature map, SIA and CIA blocks compute complementary interdependencies. We sequential arrange SIA and CIA blocks to form the IA block (see Fig. 6 (a)). IA block can be inserted at any depth of a network. Considering the computational complexity, we only place it at the bottlenecks of models where the downsampling of feature maps occurs. Multiple IA blocks located at bottlenecks of different levels can progressively enhance the feature representations with negligible number of parameters.
4 IANet for Person ReID
The architecture of IANet is illustrated in Fig. 6 (b). Here we use ResNet-50 pre-trained on ImageNet as the backbone network for person reID. The output dimension of the classification layer is set to the number of training identities. Following , we remove the last spatial down-sampling operation in the backbone network to increase retrieval accuracy with very light computation cost added. IA blocks are then inserted into the backbone network after stage-2 and stage-3 layers. The training procedure of IANet follows the standard identity classification paradigm , where the identify of each person is treated as a distinct class. IANet is end-to-end trained with cross-entropy loss. During testing, the features of probe and gallery images are extracted by IANet, and the cosine distance is used for matching.
5 Discussions
In this subsection, we give a brief discussion on the relations between our proposed IA block and some existing models.
Relations to Non-local Our IA and Non-local (NL) are both the concrete forms of self-attention. Compared to NL, IA is more suitable to reID because of the following advantages: (1) the proposed CIA is the first attempt to apply self-attention on the channel dimension, which is conductive to highlighting important but small details or body parts. (2) NL can be seen a special case of SIA in the single-context version. Multi-context SIA fuses appearance similarities across multiple patch size, which could localize the body parts more precisely. (3) SIA considers the spatial structure of pedestrians and models the location relations to constrain and complement the appearance similarity.
Relations to SCA-CNN and CBAM SCA-CNN and CBAM propose spatial and channel attention to enhance important features and suppress unnecessary ones. However there is no direct guidance for this process, making these methods easily produce unreliable attentions. On the contrary, our IA models generates the attention maps guided by semantic similarity between features, which could adaptively locate body parts and are more reliable.
Relations to Squeeze-and-Excitation CIA has some similarities with Squeeze-and-Excitation network (SE) because both are designed to model the interdependencies between channels to improve the feature representation power. However, SE computes channel-wise attention that selectively emphasizes informative features, while ignoring the spatial-wise responses due to global spatial pooling. Therefore, the spatial structure information is lost.
Relations to Graph Convolutional Network SIA and CIA could be treated as the extended Graph Convolutional Network (GCN), where the nodes of graph are defined by the spatial features and channel features, respectively, and the adjacent matrix is the semantic relation map. Compared to conventional GCN where the adjacent matrix is fixed, SIA and CIA change the graph structure adaptively during training, which is more desirable for information propagation between feature nodes.
Experiments
Datasets and Evaluation Metric. We conduct extensive experiments on four person reID benchmarks, CUHK03 , Market-1501 , DukeMTMC-reID and MSMT17 . For CUHK03, we follow the standard protocol detailed in and report the results on manually annotated and DPM-detected images. We adopt mean Average Precision (mAP) and Cumulative Matching Characteristics (CMC) as evaluation metrics.
Implementation details. For our implementation, the input images are resized to after random left-right flipping. The initial learning rate is set to . Adam optimizer is used with a mini-batch size of for training. The number of context levels ( in Eq. 2) is set to . Because the feature maps at different stages of ResNet have different spatial sizes, we use different standard deviations ( and in Eq. 3) for IA blocks at different stages. Specially, and are set to and when IA blocks are added to stage-3 layers, and and are set to and when IA blocks are added to stage-2 layers.
2 Comparison with State-of-the-art Approaches
Market-1501 and DukeMTMC. In Tab. 1, we compare IANet with state-of-the-arts on Market-1501 and DukeMTMC. IANet achieves the best performance on all evaluation criteria. It is noted that: (1) The gaps between our results and the spatial alignment methods (Spindle , AACN and PSE ) that incorporate an external part detection sub-network are significant: about improvement on top-1 accuracy and mAP. We argue that these methods are prone to performance degeneration due to inaccurate part detection. On the contrary, our method can adaptively locate the body parts guided by the semantic similarity without external network. (2) IANet outperforms the attention-centric methods (AACN* , HA-CNN and RPP ) that uses spatial attention to learn discriminative parts. We argue that these methods often fail to produce reliable attentions as there is no guidance for this process. On the contrary, our method generates spatial attention maps guided by semantic similarity between spatial features, which are more reliable. (3) IANet outperforms the multi-scale methods (DaRe , DPFL , MLFN and KPM ), with an improvement up to on mAP. The superiority of IANet over the multi-scale methods indicates that without explicitly fusing features from multiple scales, IANet could also cope with large scale changes.
CUHK03. In Tab. 2, we report the top-1 and top-5 accuracies on CUHK03. IAnet outperforms the state-of-the-arts. It is noteworthy that there is a small gap between the labeled evaluation and the detected evaluation of our method, which indicates that our method is robust at the presence of the imperfect detection.
MSMT17. We further evaluate our method on a recent large scale dataset, namely MSMT17. As shown in Tab. 3, our method significantly outperform existing works with top-1 and mAP. Since MSMT17 is the largest dataset with more than images, this result strongly demonstrates the superiority of our proposed method.
3 Ablation Study
In this section, we investigate the effectiveness of each component in IA block by conducting a series of ablation studies on Market-1501 and DukeMTMC datasets. We adopt ResNet-50 trained with cross-entropy loss as the baseline (denoted as base.). If there is no special explanation, we add the proposed blocks to the last residual block (bottleneck) of layer of ResNet-50.
Multi-context combination. We first compare the single-context interaction operations ( in Eq. 1) with different patch sizes . In this part, SIA uses single-context appearance relation map in the aggregation operation. As shown in Tab. 4a, there is an improvement in performance when is increased, showing the effectiveness of incorporating contextual information. However, as is further increased, the accuracy drops gradually. So we only use context patches with of , and in the multi-context interaction operation. We then explore three different fusion functions in the multi-context interaction operation ( in Eq. 2): element-wise maximum, summation, and product. In this part, SIA uses multi-context appearance relation map in the aggregation operation. Tab. 4b lists the comparison results of different fusion strategies. Element-wise product performs better than other functions and is therefore selected as the default fusion function.
We finally report the computation cost of SIA. For single-context SIA (=1), the relation map could be worked out by one matrix multiplication. The multi-context SIA (=1,2,3) does not incur extra multiplying operation compared to single-context, as has got the dot-product between every two spatial position. Specially, baseline requires 4.06 Multiply GFLOPs in a single forward pass for a 256128 pixel input image. Baseline added SIA requires 4.09 Multiply GFLOPS, corresponding only relative increase over ResNet50.
Effectiveness of introducing location prior. In the location relation map, the standard deviations and determine the shape of Gaussian function. In order to reduce the search space of hyper-parameters, we set to according to the input aspect ratio :. Fig. 7 shows the top-1 accuracy and mAP changes with , where SIA uses spatial location relation map in the aggregation operation. We can see that the performance of SIA is not sensitive to with a certain range of values.
We then investigate whether location relations can improve the performance of SIA. Tab. 4c compares different SIA blocks: Location, Appearance and Semantic that respectively use location (), appearance () and semantic () relation maps in aggregation operation. As seen, Location achieves better results than baseline, which demonstrates the effectiveness of introducing location relations. Besides, Semantic consistently outperforms Appearance, showing that location relations complement appearance relations, which can achieve more precise semantic relations.
Arrangement of CIA and SIA blocks. We first verify the effectiveness of CIA block in Tab. 4d by adding it to baseline. CIA outperforms baseline by in mAP, which implies that it is effective to enhance feature representation power by aggregating similar channel features. We then compare three different ways of arranging CIA and SIA blocks: parallel, sequential channel-spatial, and sequential spatial-channel. As shown in Tab. 4d, the sequential spatial-channel produces the best performance. Note that the results outperform adding CIA or SIA blocks independently, showing that utilizing both blocks is crucial and the best-arranging strategy further pushes performance.
Efficient positions to place IA blocks. Table 4e compares an IA block added to the bottlenecks of different stages of ResNet. The improvements of an IA block in stage2 and stage3 are similar, but smaller in stage1 and stage4. Therefore, we only insert IA blocks into stage2 and stage3 layers. Finally, we empirically verify that the bottlenecks are the effective positions to place IA blocks. Recent studies mainly focus on modifications within the ’convolution blocks’ rather than the ’bottlenecks’. Tab. 4f compares two different locations, where stage23 adds IA blocks to the bottlenecks, while stage23-c adds blocks to every residual block of stage2 and stage3 layers. We can clearly observe that placing the blocks at the bottlenecks is more effective. We argue that too many IA blocks are difficult to optimize on small reID datasets.
Effectiveness of IA block across different backbones. In order to further verify the validity of IA block, we try another three backbones besides the ResNet-50, i.e., ResNet32 (R32), ResNet101 (R101), and GoogleNet (GN). As shown in Tab. 5, our method (-IA) improves the performance w.r.t. different backbones consistently.
4 Visualization for Pose and Scale Robustness
To verify whether SIA can adaptively localize body parts under various poses and scales, we visualize the pixel-wise receptive fields learned with SIA. Specifically, each specific position has a corresponding sub-relative map. We define the sub-relation maps with high relation values as valid receptive fields and highlight them. As shown in Fig. 8, for each input image, we select five positions from head, torso, belt/bag, legs and shoe and show their corresponding valid receptive fields. We can clearly observe that SIA can adaptively localize the body parts and visual attributes under various poses and scales. For example, in Fig. 8 (a) and Fig. 8 (b), the person take on different poses. SIA could aggregate the features of regions corresponding to the body parts independently of the pose. In Fig. 8 (c), the person is in different scales due to the detection errors. SIA could adaptively adjust the scales of receptive fields based on the scales of body parts. In addition, the receptive fields of different body parts in SIA have different shapes and scales, which is superior to the fixed geometric receptive field in CNNs.
For CIA, it is hard to give comprehensive visualization about the relation maps directly. Instead, we show some aggregated channels to see whether they highlight small body parts or attribute areas. In Fig. 8, we display , (a) / (b, c) and channels. We find that the responses of specific body parts and attributes are noticeable after CIA enhances. For example, , and channel maps respond to attribute belt, bag and shoes. In short, the visualizations further demonstrate the necessity of modeling the interdependencies between channels for improving feature representation, especially for fine-grained attributes.
Conclusion
In this paper, we propose SIA and CIA blocks to improve the representational capacity of deep convolutional networks. SIA models the interdependencies between spatial features of convolutional feature maps. It can adaptively localize the body parts under various poses and scales. CIA models the interdependencies between channel features. It can further enhance the feature representations especially for small visual cues. Extensive experiments show that IANet outperforms state-of-the-arts on three public person reID datasets.
Acknowledgement This work is partially supported by National Key R&D Program of China (No.2017YFA0700800), Natural Science Foundation of China (NSFC): 61876171 and 61572465.