Multi-Level Factorisation Net for Person Re-Identification

Xiaobin Chang, Timothy M. Hospedales, Tao Xiang

Introduction

Person re-identification (Re-ID) aims to match people across multiple surveillance cameras with non-overlapping views. It is challenging because the visual appearance of a person across different cameras can change drastically due to many covariates such as illumination, background, camera view-angle and human pose (see Fig. 1). However, there exist identity-discriminative but view-invariant visual appearance characteristics or factors that can be exploited for person Re-ID. As illustrated in Fig. 1, such factors can be found at different semantic and abstraction levels, ranging from low-level colour and texture to high-level concepts, such as clothing type and gender. An ideal person Re-ID model should: (i) automatically learn the space of multi-level discriminative visual factors that are insensitive to viewing condition changes, and (ii) recognise and exploit them when matching testing images (as per human expert operating procedure ).

Most recent person Re-ID approaches employ deep neural networks (DNNs) to learn view-invariant discriminative features. For matching, the features are typically extracted from the very top feature layer of a trained model. A problem thus arises: A DNN comprises multiple feature extraction layers stacked one on top of each other; and it is widely acknowledged that, when progressing from the bottom to the top layers, the visual concepts captured by the feature maps tend to be more abstract and of higher semantic level. However, for Re-ID purposes, discriminative factors of multiple semantic levels should be ideally preserved in the learned features. Therefore existing Re-ID models using standard architectures have limited efficacy. These network architectures, though working very well for the object categorisation task such as ImageNet classification due to its focus on high-level semantic features, are not well-suited for the instance-level recognition task of Re-ID.

A number of recent deep Re-ID models started to model discriminative factors of multiple levels. Some focused on learning semantic visual features with additional supervision in the form of attributes . The idea is to explicitly define these factors as semantic attributes (gender, object carrying, clothing colour/texture, etc). By combining the Re-ID task with the attribute prediction task, the top layer of the model is expected to better capture these factors. However, annotating attributes is costly and error-prone; and defining exhaustively all the discriminative factors that are well-presented in data using attributes is extremely challenging. Importantly, Re-ID features are still only computed from the very top layer of a network. The others exploited the idea of multi-level fusion either in the form of attention maps computed at multiple intermediate layers of a network , or multiple body parts grouped into different levels . Nevertheless, none of these models attempted to combine discriminative feature representations computed from all layers/levels without handcrafted architecture design and/or layer selection. Furthermore, discriminative factors are either not modelled explicitly , or limited to body parts only .

In this paper, we propose a novel DNN architecture called Multi-Level Factorisation Net (MLFN) (see Fig. 2). MLFN learns identity-discriminative and view-invariant visual factors at multiple semantic levels. The overall network is composed of multiple blocks (each of which may contain multiple convolutional layers). Each block contains two components: A set of factor modules (FMs), each of which is a sub-network of identical architecture designed to model one factor, and a factor selection module (FSM) that dynamically selects which subset of FMs in the block are activated. Training this architecture results in FMs that specialise in processing different types of factors, and at different blocks represent factors of different semantic levels. For example, we find empirically that the FMs from bottom-blocks represent low-level semantic attributes such as clothing colour, and top-blocks represent high-level semantic attributes such as object carrying and gender (see Sec. 4.4.2). Importantly, the output vectors of the FSMs at different blocks provide a compact latent semantic feature at the corresponding semantic level. To benefit from combining both these multi-level semantic features and conventional deep features, MLFN concatenates the FSM output vectors of different levels into a Factor Signature (FS) feature and then fuses it with the final-layer deep feature before subjecting them to a training loss.

The MLFN architecture is noteworthy in that: (i) A compact FS is generated by concatenating the FSM output vectors from all blocks, and therefore multi-level feature fusion is obtained without exploding dimensionality; (ii) Using the FSM output vectors to predict person identity via skip connections and fusion provides deep supervision which ensures that the learned factors are identity-discriminative, but without introducing a large number of parameters required for conventional deep supervision. The proposed architecture can be interpreted in various ways: as a generalisation of ResNext , where the sub-networks within each block can be switched on and off dynamically; or as a generalisation of mixture-of-expert layers where multiple rather than solely one expert/sub-network are encouraged to be active at a time. More importantly, it extends both in that it is the selection of which factor modules or experts are active that provides a compact latent semantic feature, and enables the low-dimensional fusion across semantic levels and deep supervision.

MLFN is evaluated on three person Re-ID benchmarks, Market-1501 , CUHK03 and DukeMTMC-reID , and achieves state-of-the-art performance on all of them. Moreover, it is effective on general object categorisation tasks as demonstrated on CIFAR-100 , showing its potential beyond person Re-ID.

Related Work

Deep Neural Networks for Person Re-ID Most recent person Re-ID methods train deep DNN models with various learning objectives including classification, verification and triplet ranking losses . Once trained, these models typically extract visual features from the final layer of a network for matching. Since the feature map of each layer is used as input for the subsequent layer, it is commonly expected that the extracted features become more abstract and of higher semantic level towards the top layers . It is thus infeasible for the final layer of the network to capture discriminative visual features of all semantic levels on its own.

One approach to obtaining an appearance feature containing information from multiple semantic levels is training it to predict visual attributes . By defining and annotating diverse attributes at multiple semantic levels, and training to predict them, these models are forced to encode attribute information using their top-layer features . However, most existing methods require a manual definition of the attribute dictionary and large scale image-attribute annotation, making this approach non-scalable. In contrast, MLFN discovers discriminative latent factors with no additional supervision. Moreover, the task of learning multi-level factors is shared by all network blocks rather than burdening only the final layer.

Another approach is to complement the final layer feature with features from other layers. A couple of studies fused representations from multiple levels , but this required extra effort such as body-part detection or attention mechanisms and handcrafted architecture and/or layer selection. In contrast, MLFN parsimoniously fuses information from all levels in the deep network, which is possible because it provides a compact latent factor representation that can be easily aggregated without prohibitive feature dimensionality. Furthermore, the introduction of factor module subnets makes MLFN suitable for automated discovery of latent appearance factors. Note that orthogonal to multi-level factorisation and fusion, multi-scale Re-ID has also been studied which focuses on fusing image resolutions rather than semantic feature levels.

DNNs with Multi-Level Feature Fusion Multi-level fusion architectures have been developed in other computer vision tasks. In semantic segmentation , feature maps from selected levels are used with shortcut connections to provide multiple granularities to the segmentation output. In visual recognition, deep features from a few selected layers were merged together to improve the final-layer representation . However, features extracted from limited and manually-specified layers may not reflect the optimal choice for complementing the final representation. Very few fusion architectures on specific tasks, e.g., edge detection , fuse features from all layers/levels. These models are usually designed to have limited levels (e.g., 3∼53\sim 5), so their expressibility is limited. Our MLFN can employ very deep networks and parsimoniously fuses features from every level (block). This is because multi-level features are represented by the compact FSM output vectors rather than the original feature channels, which significantly reduces the fused feature dimensionality.

Related CNN Architectures Instead of constructing each block with holistic modules as in , a split-transform-merge strategy is used to construct the modularised block architecture in ResNeXt . A group of sub-network modules with duplicate structures are equally activated with their outputs summed up. Our MLFN leverages the ResNeXt design pattern, but extends it to include a dynamic selection of which module subset activates within each block for each image. This allows MLFN modules to specialise in processing different latent appearance factors, and the FSM output vectors to encode a compact descriptor of detected latent factors at the corresponding level.

Our MLFN architecture is also related to that of Mixture-of-Expert (MoE) models . In MoEs, a softmax activation module aims at identifying a single expert to process a given input instance. Mixture-of-Experts Layer (MoEL) models extend flat MoE to a stacked model. They have been used to separate localisation and classification tasks in a two-level MoEL model , or to implement very large neural networks by allowing each node in a cluster to run one expert in one layer of the large network . The proposed MLFN has the following key distinctions to MoE/MoEL: (1) MLFN dynamically detects multiple latent factors at each level that explain each input image jointly (e.g., a person can have both long hair and carry a bag). Thus MLFN uses sigmoid activated FSMs rather than softmax as used in MoE/MoEL, which assumes a single expert should dominate. (2) MLFN aggregates the FSM output vectors at all blocks into a Factor Signature (FS) to provide a single compact discriminative code. That is, while MoEL dynamically switches which experts process the data but otherwise only outputs the chosen experts’ opinion about the data; MLFN uses the information of which set of factor modules were chosen as a description of the data.

Our Contributions are as follows: (1) MLFN is proposed to automatically discover discriminative and view-invariant appearance factors at multiple semantic levels without auxiliary supervision. (2) A compact discriminative semantic representation (FS) is obtained by aggregating FSM output vectors at all levels of MLFN. (3) Our FS representation is complementary to the conventional deeply learned features. Using their fusion as a representation, we obtain state-of-the-art results on three large person Re-ID benchmarks.

Methodology

MLFN Architecture Our MLFN aims to automatically discover latent discriminative factors at multiple semantic levels and dynamically identify their presence in each input image. As shown in Fig. 2, NN MLFN blocks are stacked to model NN semantic levels. Let BnB_{n} denote the nnth block, n∈{1,...,N}n\in\{1,...,N\} from bottom to top. Within each BnB_{n}, there are two key components: multiple Factor Modules (FMs) and a Factor Selection Module (FSM). Each FM is a sub-network with multiple convolutional and pooling layers of its own, powerful enough to model a latent factor at the corresponding level indexed by nn. Each block BnB_{n} consists of KnK_{n} FMs with an identical network architecture. For simplicity, only one input image is considered in the following formulation. Given the image, the output of the iith, i∈{1,...,Kn}i\in\{1,...,K_{n}\} FM in BnB_{n} is denoted as

where Mn,i\bm{M}_{n,i} is a feature map with height HnH_{n}, width WnW_{n} and CnC_{n} channels.

where n∈{1,...,N}n\in\{1,...,N\}, σ(⋅)\sigma(\cdot) is an element-wise sigmoid and Anˉ\bar{\bm{A}_{n}} is the pre-activation output of the FSM.

Thus the factorised representation of an input image at the nnth level can be represented as a tuple:

Factor Signature In order to complement the final-level deep representation YN\bm{Y}_{N} (feature output of BNB_{N}) with the factorised representation learned from lower levels, a compact Factor Signature (FS) representation preserving discriminative information from all levels is computed. FS aggregates all FSM output vectors Sn\bm{S}_{n}, n∈{1,...,N}n\in\{1,...,N\}. Denoting FS as S^\hat{\bm{S}}, we have

Fusion MLFN fuses the deep features YN\bm{Y}_{N} computed from the final block BNB_{N} and the Factor Signature (FS) S^\hat{\bm{S}}. Concretely, YN\bm{Y}_{N} and S^\hat{\bm{S}} are first projected to the same feature dimension dd with projection function TT implemented as a fully connected layer. The final output representation R\bm{R} of MLFN is computed by averaging the two projected features as in Eq. 6.

Optimisation The visual appearance of each input is dynamically factorised into {Mn,Sn}\{\bm{M}_{n},\bm{S}_{n}\} at multiple semantic levels in the corresponding MLFN block Bn,n∈{1,...,N}B_{n},n\in\{1,...,N\}, as in Eq. 3. Denoting the iith FM in BnB_{n} as Fn,i(⋅)F_{n,i}(\cdot) and its parameters as θn,i\bm{\theta}_{n,i}, then

The output feature Yn\bm{Y}_{n} is computed as in Eq. 4. Assuming MLFN is subject to a final loss LL and the gradient ∂L∂Yn\frac{\partial{L}}{\partial{\bm{Y}_{n}}} can be acquired. In order to update the parameters θn,i\bm{\theta}_{n,i} in backpropagation, the following gradient is computed,

where Sn,iS_{n,i} is the FSM output corresponding to the iith FM in BnB_{n}. Combining Eq. 8 and Eq. 9, we have

where ∂L∂Yn\frac{\partial{L}}{\partial{\bm{Y}_{n}}} is back propagated from higher levels and ∂Fn,i∂θn,i\frac{\partial{F_{n,i}}}{\partial{\bm{\theta}_{n,i}}} is the gradient of an FM w.r.t its parameters. Sn,iS_{n,i} comes from the corresponding FSM. It dynamically indicates the contribution of Fn,iF_{n,i} in processing an input image.

Sn,iS_{n,i} will be close to 1 if the latent factor represented by Mn,i\bm{M}_{n,i} is identified to be present in the input. In this case, the impact of this input is fully applied on θn,i\bm{\theta}_{n,i} to adapt the corresponding FM. On the contrary, when Sn,iS_{n,i} is close to 0, it means the input only holds irrelevant or opposite latent factors to Mn,i\bm{M}_{n,i}. Therefore, the parameters in the corresponding FM are unchanged when training with this input as Sn,i≈0S_{n,i}\approx 0 stops the update.

The factor selection vectors Sn\bm{S}_{n} (Eq. 2) play a key role in MLFN during both training (as analysed above) and inference (providing the factor signature). Learning discriminative FSMs would be hard if trained with gradients back propagated through many blocks from the top. This is because the supervision from the loss would be indirect and weak for the FSMs at the bottom levels. However, because our final feature output R\bm{R} is computed by fusing the final-block output YN\bm{Y}_{N} with the FS S^\hat{\bm{S}} (Eq. 6), and the FS is generated by concatenating all FSM output vectors, the supervision flows from the loss down to every FSM via direct shortcut connections (Fig. 2). Thus our FSMs are deeply supervised to ensure that they are discriminative, but without the increase in parameters that would be required for deep supervision of conventional deep features.

MLFN for Person Re-ID The training procedure of MLFN for Person Re-ID follows the standard identity classification paradigm where each person’s identity is treated as a distinct class for recognition. A final fully connected layer is added above the representation R\bm{R} that projects it to a dimension matching the number of training classes (identities), and the cross-entropy loss is used. MLFN is then end-to-end trained. It discovers latent factors with no supervision other than person identity labels for the final classification loss. During testing, appearance representations R\bm{R} (Eq. 6) are extracted from gallery and probe images, and the L2 distance is used for matching.

Experiments

Datasets Three person Re-ID benchmarks, Market-1501 , CUHK03 and DukeMTMC-reID are used for evaluation. Market-1501 has 12,936 training and 19,732 testing images with 1,501 identities in total from 6 cameras. Deformable Part Model (DPM) is used as the person detector. We follow the standard training and evaluation protocols in where 751 identities are used for training and the remaining 750 for testing. CUHK03 consists of 13,164 images of 1,467 people. Both manually labelled and DPM detected person bounding boxes are provided. We adopt two experimental settings on this dataset. The first setting, denoted as CUHK03 Setting 1, is the 20 random train/test splits used in which selects 100 identities for testing and training with the rest. Results on the more challenging yet more realistic detected person bounding boxes are reported under this setting. The other setting, denoted as CUHK03 Setting 2, was proposed in . It is more challenging than Setting 1 with less training data. In particular, 767 identities are used for training and the remaining 700 identities for testing. DukeMTMC-reID is the Person Re-ID subset of the Duke Dataset . There are 16,522 training images of 702 identities, 2,228 query images and 17,661 gallery images of the other 702 identities. Manually labelled pedestrian bounding boxes are provided. Our experimental protocol follows that of . In addition to the Re-ID datasets, an object category classification dataset, CIFAR-100 , is used to show that our MLFN can also be applied to other general recognition problems. CIFAR-100 has 60K images with 100 classes with 600 images in each class. 50K images are used for training and the remaining for testing.

Evaluation metrics We use the Cumulated Matching Characteristics (CMC) curve to evaluate the performance of Re-ID methods. Due to space limitation and for easier comparison with published results, we only report the cumulated matching accuracy at selected ranks in tables rather than plotting the actual curves. Note that we also use mean average precision (mAP) as suggested in to evaluate the performance. For CIFAR100, the error rate is used.

MLFN Architecture Details For Person Re-ID tasks, sixteen blocks (N=16N=16) are stacked in MLFN. Within each building block, 32 FMs are aggregated as in . Correspondingly, a 32-D FSM output vector is generated within each MLFN block. As a result, the FS dimension K=512K=512 (3232 FMs ×16\times 16 blocks). The final feature dimension of R\bm{R}, dd is set to 1024. For the object categorisation task CIFAR-100 , we reduced the MLFN depth in order to fit the memory limitation of a single GPU. The number of blocks is reduced to 9 which results in K=288K=288. More discussion on parameter selection can be found in the Supplementary Material.

Data Augmentation The input image size is fixed to 256×128256\times 128 for all person Re-ID experiments. Left-right flip augmentation is used during training. For CIFAR-100, training images are augmented as in . No data augmentation is used for testing.

Optimisation Settings All person Re-ID models are fine-tuned on ImageNet pre-trained networks. The Adam optimiser is used with a mini-batch size of 64. Initial learning rate is set to 0.00035 for all Re-ID datasets except CUHK03 setting 2 with 0.0005. Similarly, Training iterations are 100k for all Re-ID datasets except CUHK03 setting 2 for which it is 75k. For CIFAR, the initial learning rate is set to 0.1 with a decay factor 0.1 at every 100 epochs and Nesterov momentum of 0.9. SGD optimisation is used with a 256 mini-batch size on a K80 GPU for 307 epochs training.

2 Person Re-ID Results

Results on Market-1501 Comparisons between MLFN and 14 state-of-the-art methods on Market-1501 are shown in Table 1. SQ and MQ correspond to the single and multiple query setting respectively . The results show that our MLFN achieves the best performance on all evaluation criteria under both settings. It is noted that: (1) The gaps between our results and those of the two models that attempt to fuse multi-level features are significant: 13.1% R1 accuracy improvement under SQ. This suggests that our fusion architecture with deep supervision is more effective than the handcrafted architectures with manual layer selection in , which require extra effort but may lead to suboptimal solutions. (2) The best model that uses attribute annotation also yields inferior results (SQ 83.6 vs 90.0 for R1 and 62.6 vs 74.3 for mAP), despite the fact that more supervision was used. This indicates that the automatically discovered latent factors at multiple levels in MLFN provides a more discriminative representation. (3) The closest competitor, DPFL uses multiple network branches to model image input scaled to different resolutions, which is orthogonal to our approach and can be easily combined to improve our performance further.

Results on CUHK03 Table 2 shows results on CUHK03 Setting 1 when detected person bounding boxes are used for both training and testing. MLFN achieves the best result, 82.8%, under this setting. Note that DGD , Spindle Net and HP-net were trained with the JSTL setting where additional data in the form of six Re-ID datasets were used. They also used mixed labelled and detected bounding boxes for both training and test. Following the multi-bounding box setting, even without using auxiliary training data as in JSTL, the accuracy of MLFN jumps from 82.8% to 89.2%. Similarly, LSRO used external Re-ID datasets for training, thus gaining an advantage.

The results in Table 3 correspond to CUHK03 Setting 2, which is a harder and newer setting with less reported results. Clear gaps are now shown between MLFN and DPFL : The rank 1 (R1) performance of MLFN is more than 11% higher using either labelled or detected person images. This result suggests that the advantage of MLFN is more pronounced given less training data. Similar performance jumps are also observed using the mAP metric.

Results on DukeMTMC-reID Person Re-ID results on DukeMTMC-reID are given in Table 4. This dataset is challenging because the person bounding box size varies drastically across different camera views, which naturally suits the multi-scale Re-ID models such as DPFL . The results show that MLFN is 1.8% and 2.2% higher than the prior state-of-the-art DPFL on R1 and mAP metrics respectively. This indicates that even without explicitly extracting features from input images scaled to different resolutions, by fusing features from multiple levels (blocks in MLFN), it can cope with large scale changes to some extent.

3 Object Categorisation Results

We next evaluate whether our MLFN is applicable to more general object categorisation tasks by experimenting on CIFAR-100. The results are shown in Table 5. For direct comparison we reproduce results with ResNet and ResNeXt of similar depth and model size to our MLFNThe results in Table 5 still have a gap to the state-of-the-art results such as . The latter were obtained with much larger networks. Those models and batch sizes are beyond the GPU resources at our disposal. . The improved result over ResNeXt shows that our dynamic factor module selection and factor signature feature bring clear benefit. MLFN also beats DualNet , another representative recent ResNet-based model that fuses two complementary ResNet branches as in an ensemble, thus doubling in model size. Note that for distinguishing different object categories, e.g., dog and bird, low-level factors such as colour and texture are often less useful as for instance classification problems such as person Re-ID. However, this result suggests that discriminative latent factors still exist in multiple levels for object categorisation and can be discovered and exploited by our MLFN.

4 Further Analysis

Recall that our MLFN discovers multiple discriminative latent factors at each semantic level, by aggregating FMs with identical structures within each block BnB_{n}. The FSM output vectors Sn\bm{S}_{n} enable dynamic factorisation of an input image into distinctive latent attributes, and these are aggregated over all blocks into a compact FS feature (S^\hat{\bm{S}}) for fusion (Eq. 6) with the conventional (final-block) deep feature YN\bm{Y}_{N} to produce the final representation R\bm{R}. To validate the contributions of each component, we compare: MLFN: Full model. MLFN-Fusion: MLFN using dynamic factor selection, but without fusion of the FS feature. ResNeXt: When the FSMs are removed so all FMs are always active, our model becomes ResNeXt . ResNet: When the sub-networks at each level of ResNeXt are replaced with one larger holistic residual module, we obtain ResNet .

A comparison of these models on all three person Re-ID datasets is shown in Table 6. We can see that MLFN is consistently better than the stripped-down versions on all datasets, and each new component contributed to the final performance: The margin between MLFN and MLFN-Fusion shows the importance of including the latent factor descriptor FS in the person representation and suggests that the FS feature is complement to the final-block feature YN\bm{Y}_{N}, and the margin between MLFN−-Fusion and ResNeXt shows the benefit of dynamic module selection.

4.2 Analysis on Latent Factors

Recall that a key idea of MLFN is to extract the factor signature S^\hat{\bm{S}} by aggregating FSM outputs of all blocks and using it as

Efficacy of Re-ID with Factor Signature Alone For solely FS-based matching, we train a binary SVM based on the absolute difference of paired FS to predict whether they belong to the same person or not. SVM scores of testing pairs are then computed for recognition. The corresponding results on Market-1501 are reported in Table 7. It shows that, compared with the results in Table 1, the result of FS only is already comparable with the state-of-the-art.

Discovered Latent Factors are Predictive of Attributes What do the discovered latent factors represent? We hypothesise that despite not being trained with any manually annotated attributes, FS (S^\hat{\bm{S}}) is identifying latent data-driven attributes present in the data; these latent attribute may overlap or correlate with human-defined semantic attributes. To validate this, SVMs are then trained based on S^\hat{\bm{S}} only to predict ground-truth manually annotated attributes in Market-1051 and DukeMTMC-reID. Results based on the final representation R\bm{R} from MLFN are also reported. Finally, these are compared to APR , which is end-to-end trained based on attribute supervision.

On Market-1051, MLFN-S^\hat{\bm{S}} and APR obtain the same performance of 85.33%85.33\%. MLFN-R\bm{R} further improves to 87.50%87.50\%. On DukeMTMC-reID, 82.30%82.30\% and 83.58%83.58\% are achieved by MLFN-S^\hat{\bm{S}} and MLFN-R\bm{R} respectively, which are better than APR’s 80.12%80.12\%. These results thus show that our low-dimensional MLFN-S^\hat{\bm{S}} alone can be more effective in attribution prediction than APR. Remind that MLFN is trained without annotated attributes while APR network is designed for supervised attribute learning. This shows that our architecture is well suited for extracting semantic attribute related information automatically. More analysis of the relations between FS and Attributes can be found in the Supplementary Material.

What is Learned To visualise the latent discriminative appearance factors learned by MLFN, we rank each element of the FSM output vector, denoted as Sn,iS_{n,i}, with all testing samples in Market-1501 as inputs. Person images with the highest and lowest twenty values of each Sn,iS_{n,i} are recorded. Figure 3 shows four example sets of such images from different element i,i∈{1,..,Kn}i,i\in\{1,..,K_{n}\} and blocks n,n∈{1,...,N}n,n\in\{1,...,N\}. Clear visual semantics can be seen from both the highest and lowest FSM output value image clusters in each group. And as expected, as the block index number nn increases, the semantic level of the latent factors captured at the corresponding blocks gets higher, i.e., they evolve from colour and texture related factors to clothes style and gender related ones. This is achieved despite that no attribute supervision is used in training MLFN. It is also interesting to note that visual characteristics conveyed by images with the highest FSM output values are complementary or opposite to those of lowest ones from the same group. For example, highest value images in S2,29S_{2,29} contain green colour, while lowest value images contain the complementary colour red. High value in S7,31S_{7,31} encodes cold colours while low value encodes warm colours. Highest values in S10,23S_{10,23} reflect textures while lowest ones mean large untextured colour blocks are detected. Images of men select with high confidence S15,29S_{15,29}, while images of females depress its value.

Conclusion

We proposed MLFN, a novel CNN architecture that learns to discover and dynamically identify discriminative latent factors in input images for person Re-ID. The factors computed at different levels of the network correspond to latent attributes of different semantic levels. When the selections of the factors are used as a feature and fused with the conventional deep feature, a powerful view-invariant person representation is obtained. MLFN obtains state-of-the-art results on three largest Re-ID datasets, and shows promising results on a more general object categorisation task.

Appendix A Supplementary Material

The number of blocks (NN) in MLFN is set to 1616 follows the ResNeXt-50 architecture. The FS dimension KK depends on NN and the number of FMs at each MLFN block. We set these, without tuning, so that the model is of a comparable overall size to ResNeXt-50 for direct comparison. On our GTX1080 GPU, the runtime is similar: MLFN (0.810.81s/batch) and ResNeXt (0.780.78s/batch), and so is the GPU memory consumption. The final feature dimension dd of MLFN is set to 10241024 since it is the widely used feature dimension for Person ReID such as . The impacts of different dd values on the re-id performance are illustrated as in Figure 4. It can be seen that the performance is consistently good when d>512d>512.

The proposed MLFN architecture consists of 16 MLFN Blocks/Layers. A Factor Selection Module (FSM) is included in each Block. The FSM networks used in this paper are all three-layered Multiple Layer Perceptron (MLP). Global Average Pooling (GAP) is applied on the input of FSM. Batch Normalisation and Relu are used to activate each layer’s output. Architecture details are shown in Table 8.

A.2 Examples of FS Predicted Attributes

In Sec. 4.4.2 of the main paper, we have shown that the attribute prediction accuracy obtained with the factor signature (FS, S^\hat{\bm{S}}) alone in the proposed MLFN is already better than a supervised attribute prediction model APR (e.g., 82.30% vs 80.12% on DukeMTMC-reID). Here, we show some qualitative results.

Figure 5 shows three examples where the predicted attributes using our FS feature and the human labelled attributes are compared. For each person image, 35 binary attributes are annotated by human annotators on the identity level, that is, different images of the same person would have the identical attribute vectors regardless whether those attributes are visually observable in the images. These attributes form different groups and within each group, they are mutually exclusive. For example, female and male form one group, and young, teen, adult, old form another. Some attributes are thus subjective, e.g., no ground-truth age is known and there is no clear definition of what ‘young’ entails.

Figure 5(a) shows an example where our FS feature can be used to correctly predict all the attributes with SVM classifiers. In this example, although the big hat occludes the face and part of the hair of the person, the colour of the top and the shoe style give away the fact that this a female. A harder example is shown in Figure 5(c). This time the image is a bit blurred and the viewpoint is from the back. However, our FS feature can still predict all the attributes correctly. Our FS feature based prediction makes two mistakes for the person image shown in Figure 5(e). Specifically, the backpack attribute is missed and the lower-body garment colour is predicted to be black rather than blue. Both mistakes are understandable. For the backpack, since the frontal view is shown and the backpack has very thin straps, this attribute can be easily missed even by human (the human annotator labelled this because s/he had access to multiple views of this person including a back view where the backpack is clearly visible). As for the blue vs black for the lower-body cloth, it seems to be a close call even for humans.

References