Fashion Landmark Detection in the Wild

Ziwei Liu, Sijie Yan, Ping Luo, Xiaogang Wang, Xiaoou Tang

Introduction

Visual fashion analysis has drawn lots of attentions recently, due to its wide spectrum of applications such as clothes recognition , retrieval , and recommendation . It is a challenging task because of the large variations presented in the clothing items, such as the changes of poses, scales, and appearances. To reduce these variations, existing works tackled the problem by looking for informative regions, i.e. detecting the clothes bounding boxes or the human joints . We go beyond the above by studying a more discriminative representation, fashion landmark, which is the key-point located at the functional region of clothes, for example the neckline and the cuff.

This work addresses fashion landmark detection or fashion alignment in the wild. Different from human pose estimation, which detects human joints such as neck and elbows as shown in Fig.1 (a.1), fashion alignment localizes fashion landmarks as shown in (a.2). These landmarks facilitate fashion analysis in the sense that they not only implicitly capture bounding boxes of clothes, but also indicate their functional regions, which can better distinguish design/pattern/category of the clothes. Therefore, features extracted from these landmarks are more discriminative than those extracted from human joints. For example, when search for a dress with ‘V-neck and fringed-hem’, it is more desirable to extract features from collar and hemline.

To fully benchmark the task of fashion landmark detection, we select a large subset of images from the DeepFashion database to constitute a fashion landmark dataset (FLD). These images have large pose and scale variations. With FLD, we show that fashion landmark detection in clothes images is a more challenging task than human joint detection in three aspects. First, clothes undergo non-rigid deformations or scale variations as shown in Fig.1 (a.3-4), while rigid deformations are usually presented in human joints. Second, fashion landmarks exhibit much larger spatial variances than human joints, as illustrated in Fig.1 (b), where we plot the positions of the landmarks and the relative human joints in the test set of the FLD dataset. For instance, the positions of ‘left sleeve’ are more diverse than those of the ‘left elbow’ in both the vertical and horizontal directions. Third, the local regions of fashion landmarks also have larger appearance variances than those of human joints. As shown in Fig.1 (c), we average the patches centered at the fashion landmarks and human joints respectively, resulting in several visual comparisons. The patterns of the mean patches of human joints are still recognizable, but those of the mean patches of fashion landmarks are not.

To tackle the above challenges, we propose a deep fashion alignment (DFA) framework, which cascades three deep convolutional networks (CNNs) for landmark estimation. It has three appealing properties. First, to ensure the CNNs have high discriminative powers, unlike existing works that only estimated the landmarks’ positions, we train the cascaded CNNs to predict both the landmarks’ positions and the pseudo-labels, which encode the similarities between training samples to boost the estimation accuracy. In each stage of the network cascade, the scheme of pseudo-label is carefully designed to reduce different variations presented in the fashion images. Second, instead of training multiple networks for each body part as previous work did , the DFA framework trains CNNs using the full image as input to significantly reduce computations. Third, in the DFA framework, an auto-routing strategy is introduced to partition the challenging and easy samples, such that different samples can be handled by different branches of CNNs.

Extensive experiments demonstrate the effectiveness of the proposed method, as well as its generalization ability to pose estimation. Fashion landmark is also compared to clothing bounding boxes and human joints in two applications, fashion attribute prediction and clothes retrieval, showing that fashion landmark is a more discriminative representation to understand fashion images.

Visual Fashion Understanding Visual fashion understanding has been a long-pursuing topic due to the many human-centric applications it enables. Recent advances include predicting semantic attributes , clothes recognition and retrieval and fashion trends discovery . To better capture discriminative information in fashion items, previous works have explored the usage of full image , general object proposals , bounding boxes and even masks . However, these representations either lack sufficient discriminative ability or are too expensive to obtain. To overcome these drawbacks, we introduce the problem of clothes alignment in this work, which is a necessary step toward robust fashion recognition.

Human Pose Estimation We further convert the problem of clothes alignment into fashion landmark estimation. Though there are no prior work in fashion landmarks, approaches from similar fields (e.g. human pose estimation ) serve as good candidates to explore. Recently, deep learning has shown great advantages in locating human joints and there are generally two directions here. The first direction utilizes the power of cascading for iterative error correction. DeepPose employs devide-and-conquer strategy and designs a cascaded deep regression framework on part level while Iterative Error Feedback emphasizes more on the target scheduling in each stage. The second direction , on the other hand, focuses on the explicit modeling of landmark relationships using graphical models. Chen et al. proposed the combination of CNN and structural SVM to model the tree-like relationships among landmarks while Thompson et al. plugged Markov Random Field (MRF) into CNN for joint training. Here, our framework attempts to absorb the advantages of both directions. The cascading and auto-routing mechanisms enable both stage-wise and branch-wise variation reductions while pseudo-labels encode multi-level sample relationships which depict typical global and local landmark configurations.

Fashion Landmark Dataset (FLD)

To benchmark fashion landmark detection, we select a subset of images with large pose and scale variations from DeepFashion database , to constitute FLD. We label and refine landmark annotations in FLD to make sure each image is correctly labeled with 88 fashion landmarks along with their visibilityThree states of visibility are defined for each landmark, including visible (located inside of the image and visible), invisible (inside of the image but occluded), and truncated/cut-off (outside of the image).. Overall, FLD contains more than 120K120K images. Sample images and annotations are shown in Fig.2 (a). To characterize the properties of FLD, we divide the images into five subsets according to the positions and visibility of their ground truth landmarks, including the subsets of normal/medium/large poses and medium/large zoom-ins. The ‘normal’ subset represents images with frontal pose and no cut-off landmarks. The subsets of ‘medium’ and ‘large’ poses contain images with side or back views, while the subsets of ‘medium’ and ‘large’ zoom-ins contain clothing items with more than one or three cut-off landmarks, respectively. Sample images of the five subsets are illustrated in Fig.2 (b) and their statistics are demonstrated in Fig.2 (c), which shows that FLD contains substantial percentages of images with large poses/scales.

Our Approach

Fashion landmark exhibits large variations in both spatial and appearance domain (see Fig.1 (b)(c)). Fig.2 (c) further shows that more than 30%30\% images have large pose and zoom-in variations. To account for these challenges, we propose a deep fashion alignment (DFA) framework as shown in Fig.3 (a), which consists of three stages, where each stage subsequently refines previous predictions. Unlike the existing representative model for human joints prediction, such as DeepPose as shown in (b), which trained multiple networks for all localized parts in each stage, the proposed DFA framework that functions on full image is able to achieve superior performance with much lower computations.

Framework Overview As shown in Fig.3 (a), DFA has three stages. In each stage, VGG-16 is employed as network architecture. In the first stage, DFA takes the raw image II as input and predicts rough landmark positions, denoted as l1^\hat{l^{1}}, as well as pseudo-labels, denoted as f1^\hat{f^{1}}, which represent landmark configurations such as clothing categories and poses. In the second stage, both the input image II and the predictions of stage-1, l1^\hat{l^{1}}, are fed in. The whole network is required to predict landmark offsets, signified as δl2^\hat{\delta l^{2}}, and pseudo-labels f2^\hat{f^{2}} that represent local landmark offsets. The landmark prediction of stage-2 is computed as l2^=l1^+δl2^\hat{l^{2}}=\hat{l^{1}}+\hat{\delta l^{2}}. The third stage has two CNNs as two branches, which have identical input and output. Similar to the second stage, each CNN employs image II as input and learns to estimate landmark offsets δl3^\hat{\delta l^{3}} and pseudo-labels f3^\hat{f^{3}}, which contains information about contextual landmark offsets. In stage-3, each image is passed through one of these two branches. The selection of branch is determined by the predicted pseudo-labels f2^\hat{f^{2}} in stage-2. The final prediction is computed as l3^=l2^+δl3^\hat{l^{3}}=\hat{l^{2}}+\hat{\delta l^{3}}.

Network Cascade Cascade has been proven an effective technique for reducing variations sequentially in pose estimation . Here, we build the DFA system by cascading three CNNs. The first CNN directly regresses landmark positions and visibility from input images, aided by pseudo-labels in the space of landmark configuration. These pseudo-labels are achieved by clustering absolute landmark positions, indicating different clothing categories and poses, as shown in Fig.4 stage 1. For example, cluster 1 and 4 represents ‘short T-shirt in front view’ and ‘long coat in side view’, respectively. The second CNN takes both the input image and the predictions from stage-1 as input, and estimates the offsets that should be made to correct the predictions of the first stage. In this case, we are learning in the space of landmark offset. Thus, the pseudo-labels generated here represent typical error patterns and their magnitudes, as shown in Fig.4 stage 2. For instance, cluster 1 represents the corrections should be made along upside and downside, while cluster 2 suggests left/right-direction corrections. In stage-3, we partition data into two subsets, according to the error patterns predicted in the second stage as shown in Fig.4 stage 3, where branch one deals with ‘easy’ samples such as frontal T-shirt with different sleeve length, while branch two accounts for ‘hard’ samples such as selfie pose (cluster 1), back view (cluster 2) and large zoom-in (cluster 3).

where dist(⋅)dist(\cdot) is a distance measure. CkC_{k} is a set of kk cluster centers, obtaining by k-means algorithm on the spaces of landmark coordinates (stage-1) or offsets (stage-2 and 3). TT is the temperature parameter to soften pseudo-labels. We adopt K=20K=20 for all three stages.

Here we explain the pseudo-label in each of the three stages. Cluster centers Ck1C^{1}_{k} in stage-1 are obtained in landmark configuration space lil_{i}, where lil_{i} is the ground truth landmark positions for sample ii. Then pseudo-label fi1f^{1}_{i} of sample ii in stage-1 can be written as fi1(k)=exp⁡(−∥li−Ck1∥2T)f^{1}_{i}\left(k\right)=\exp\left(-\frac{\|l_{i}-C^{1}_{k}\|_{2}}{T}\right). We now have a landmark estimation li1^\hat{l^{1}_{i}} from stage-1 for sample ii. In stage-2, we first define the landmark offset δli2=li1^−li\delta l^{2}_{i}=\hat{l^{1}_{i}}-l_{i}, which is the correction should be made on stage-1 estimation. Cluster centers Ck2C^{2}_{k} in stage-2 are obtained in landmark offset space δli2\delta l^{2}_{i}. Similarly, pseudo-label fi2f^{2}_{i} of sample ii in stage-2 can be written as fi2(k)=exp⁡(−∥δli2−Ck2∥2T)f^{2}_{i}\left(k\right)=\exp\left(-\frac{\|\delta l^{2}_{i}-C^{2}_{k}\|_{2}}{T}\right). Since outer product ⊗\otimes of two offsets contains the correlations between different fashion landmarks (e.g. ‘left collar’ v.s. ‘left sleeve’), we further include these contextual information into the pseudo-labels of stage-3 f3f^{3}. To make the results of outer product comparable, we convert them into vectors by stacking columns, which is denoted as lin()lin\left(\right). The landmark offset of stage-3 is defined as δli3=li2^−li\delta l^{3}_{i}=\hat{l^{2}_{i}}-l_{i}, where li2^=li2^+δli2^\hat{l^{2}_{i}}=\hat{l^{2}_{i}}+\hat{\delta l^{2}_{i}} is the estimation made by stage-2. Thus, cluster centers Ck3C^{3}_{k} in stage-3 are obtained in contextual offset space δcontextli3=lin(δli3⊗δli3)\delta_{context}l^{3}_{i}=lin\left(\delta l^{3}_{i}\otimes\delta l^{3}_{i}\right). Similarly, pseudo-label fi3f^{3}_{i} of sample ii in stage-3 can be written as fi3(k)=exp⁡(−∥δcontextli3−Ck3∥2T)f^{3}_{i}\left(k\right)=\exp\left(-\frac{\|\delta_{context}l^{3}_{i}-C^{3}_{k}\|_{2}}{T}\right). The pseudo-labels used in each stage are summarized in Table 1.

Auto-Routing Another important building block of DFA is the auto-routing mechanism. It is built upon the fact that the estimated pseudo-labels in stage-2 f2^\hat{f^{2}} reflects the error patterns for each sample. We first associate each cluster center with an average error magnitude e(Ck2),k=1…Ke\left(C^{2}_{k}\right),k=1\ldots K. This can be done by averaging the errors of training samples in each cluster. Then, we define the error function G(⋅)G\left(\cdot\right) within pseudo-label fi2^\hat{f^{2}_{i}} for sample ii in stage-2: G(fi2^)=∑k=1Ke(Ck2)⋅fi2^(k)G\left(\hat{f^{2}_{i}}\right)=\sum_{k=1}^{K}e\left(C^{2}_{k}\right)\cdot\hat{f^{2}_{i}}\left(k\right). Therefore, the routing function rir_{i} for sample ii is formulated as

where 1(⋅)\textbf{1}\left(\cdot\right) is the indicator function and ϵ\epsilon is the error threshold for auto-routing. We set ϵ=0.3\epsilon=0.3 empirically. If ri=1r_{i}=1 indicates sample ii will go through branch 11 in stage-3, and ri=0r_{i}=0 indicates otherwise.

Training Each stage of DFA is trained with multiple loss functions, including landmark estimation LpositionsL_{positions}, visibility prediction LvisibilityL_{visibility}, and pseudo-label approximation LlabelsL_{labels}. The overall loss function LoverallL_{overall} is

where l^\hat{l}, v^\hat{v} and f^\hat{f} are the predicted landmark positions, visibility, and pseudo-labels respectively. We employ the Euclidean loss for LpositionsL_{positions} and LlabelsL_{labels}, and the multinomial logistic loss is adopted for LvisibilityL_{visibility}. α(t)\alpha\left(t\right) and β(t)\beta\left(t\right) are the balancing weights between them. All the VGG-16 networks are pre-trained using ImageNet and the entire DFA cascaded network is fine-tuned by stochastic gradient decent with back-propagation.

The proper scheduling of α(t)\alpha\left(t\right) and β(t)\beta\left(t\right) is very important for network performance. If they are too large, it disturbs the training of landmark positions. If they are too small, the training procedure cannot benefit from these auxiliary information. Similar to , we design a piecewise adjustment strategy for α(t)\alpha\left(t\right) and β(t)\beta\left(t\right) during training process,

where t1=2000t_{1}=2000 iterations and t2=4000t_{2}=4000 iterations in our implementation. The adjustment for β(t)\beta\left(t\right) takes a similar form.

Computations For a three stage cascade to predict 88 fashion landmarks, DeepPose is required to train 1717 VGG-16 models in total, while only three VGG-16 models need to be trained for DFA. Thus, our proposed approach at least saved 55 times computational costs.

Experiments

This section presents evaluation and analytical results of fashion landmark detection, as well as two applications including clothing attribute prediction and clothes retrieval.

Competing Methods Since this work is the first study of fashion landmark detection, it is difficult to find direct comparisons. Nevertheless, to fully demonstrate the effectiveness of DFA, we compare it with two deep models, including DeepPose and Image Dependent Pairwise Relations (IDPR) , which achieved best-performing results in human pose estimation. They are two representative methods that explored network cascade and graphical model to handle human pose. Specifically, DeepPose designed a cascaded deep regression framework on human body parts, while IDPR combined CNN and structural SVM to model the relations among landmarks. To have a fair comparison, we replace the backbone networks in DeepPose and IDPR with VGG-16 and carefully train them using the same data and protocol as DFA did.

We demonstrate the merits of each component in DFA.

Effectiveness of Network Cascade Table 2 lists the performance of NE among three stages, where we have two observations. First, as shown in stage one, training DFA with both landmark position and visibility, denoted as ‘direct regression’, outperforms training DFA with only landmark position, denoted as ‘- visibility’, showing that visibility helps landmark detection because it indicates variations of pose and scale (e.g. zoom-in). Second, cascade networks gradually reduce localization errors on all fashion landmarks from stage one to stage three. By predicting the corrections over previous stage, DFA decomposes a complex mapping problem into several subspace regression. For example, Fig. 8 (e-g) demonstrates the quantitative stage-wise landmark detection results of DFA on different clothing items. In these cases, stage-1 gives rough predictions with shape constraints, while stage-2 and stage-3 refine results by referring to local and contextual correction patterns.

Effectiveness of Different Pseudo-Labels Within each stage, choices of different pseudo-labels are explored. We design pseudo-labels representing landmark configurations, local landmark offsets and contextual landmark offsets for three stages respectively. Table 3 shows that pseudo-labels lead to substantial gains beyond direction regression, especially for the first stage. Next, we further justify the forms of different pseudo-labels adopted for each stage. In stage-1, we find that using soft label, denoted as ‘+p. labels (T=20)+p.~{}labels~{}(T=20)’, instead of hard label, denoted as ‘+p. labels (T=1)+p.~{}labels~{}(T=1)’, results in better performance, because soft label is more informative in identifying landmark configuration of samples. In stage-2, pseudo-label generated from offset landmark positions is superior to that generated from absolute landmark positions since landmark offsets can provide more guidance on the local corrections to be predicted. In stage-3, including contextual landmark offsets help achieve further gains, due to the fact that landmark corrections to be made are generally correlated.

Effectiveness of Auto-Routing Finally, we show that auto-routing is an effective way to tackle data with different correction difficulties. From Table 2 stage-3, we can see that auto-routing (denoted as ‘+auto-routing’) provides more benefits when compared with averaging the predictions from two branches trained using all data (denoted as ‘+two-branch’). By further inspecting the stage-wise performance on each evaluation set, which is shown in Table 3, we can observe that auto-routing mechanism improves the performance of medium/large zoom-in subsets, showing that the routing function makes one of the branch in stage-3 focus on difficult samples.

2 Benchmarking

To illustrate the effectiveness of DFA, we compare it with state-of-the-art human pose estimation methods like DeepPose and IDPR . We also analyze the strengths and weaknesses of each method on fashion landmark detection.

Landmark Types Fig. 5 (the first row) shows the percentage of detection rates on different fashion landmarks, where we have three observations. First, on landmark ‘hem’, DeepPose performs better when distance threshold is small, while IDPR catches up when the threshold is large, because DeepPose is a part-based method which can locally refines the results for easy cases. Second, collars are the easiest landmarks to be detected while sleeves are the hardest. Third, DFA consistently outperforms both DeepPose and IDPR or shows comparable results on all fashion landmarks, showing that the pseudo-labels and auto-routing mechanisms enable robust fashion landmark detection.

Clothing Types Fig. 5 (the second row) shows the percentage of detection rates on different clothing types. Again, DFA outperforms all other methods on all distance thresholds. We have two additional observations. First, DFA (stage-1) already achieves comparable results on full-body and lower-body clothes when compared with IDPR and DeepPose (stage-3). Second, upper-body clothes pose most challenges on fashion landmark detection. It is partially due to the various clothing sub-categories contained.

Difficulty Levels Fig.6 shows the percentage of detection rates on different evaluation subsets, with the distance threshold fixed at 15 pixels. Two observations are made here. First, fashion landmark detection is a challenging task. Even the detection rate for normal pose set is just above 70%70\%. More powerful model needs to be developed. Second, DFA has the most advantages on medium pose/zoom-in subsets. Pseudo-labels provide effective shape constraints for hard cases. Please also note that DFA (stage-3) requires much less computational costs than DeepPose (stage-3).

For a 300×300300\times 300 image, DFA takes around 100ms100ms to detect full sets of fashion landmarks on a single GTX Titan X GPU. In contrast, DeepPose needs nearly 650ms650ms in the same setting. Our framework has large potential in real-life applications. Visual results of fashion landmark detection by different methods are given in Fig.8.

3 Generalization of DFA

To test the generalization ability of the proposed framework, we further apply DFA on a related task, i.e. human pose estimation, as reported in Table 4. In the following, DFA is trained and evaluated on LSP dataset as did.

First, we compare DFA system with other state-of-the-art methods on pose estimation task. Without much adaptation, DFA achieves 74.474.4 mean strict PCP results, with 8787, 9191, 7070, 5656, 8181, 7676 for ‘torso’, ‘head’, ‘u.arms’, ‘l.arms’, ‘u.legs’ and ‘l.legs’ respectively. It shows comparable results to and outperforms several recent works , showing that DFA is a general approach for structural prediction problem besides fashion landmark detection.

Then, we show that pseudo-label and auto-routing scheme of DFA can be generalized to improve performance of pose estimation methods, such as IDPR . trained DCNN and achieved 7575 mean strict PCP. We add pseudo-labels to this DCNN and include auto-routing in cascading predictions. Training and evaluation of graphical model are kept unchanged. Pseudo-labels leverage the result to 7777 mean strict PCP and auto-routing leads to another 1.61.6 point gain. It demonstrates that pseudo-labels and auto-routing are effective and complementary techniques to current methods.

4 Applications

Finally, we show that fashion landmarks can facilitate clothing attribute prediction and clothes retrieval. We employ a subset of DeepFashion dataset , which contains 10K10K images, 5050 clothing attributes and corresponding image pairs (i.e. the images containing the same clothing item). We compare fashion landmarks with different localization schemes, including the full image, the bounding box (bbox) of clothing item, and the human-body joints, where fashion landmarks are detected by DFA, human joints are obtained by the executable code of , and bounding boxes are manually annotated. For both tasks of attribute recognition and clothes retrieval, we use off-the-shelf CNN features as described in .

Attribute Prediction We train a multi-layer perceptron (MLP) using the extracted CNN features as input to predict all 5050 attributes. Following , we employ the top-kk recall rate as measuring criteria, which is obtained by ranking the classification scores and determine how many ground truth attributes have been found in the top-kk predicted attributes. Overall, the average top-55 recall rates on 50 attributes of ‘full image’, ‘bbox’, ‘human joints’, and ‘fashion landmarks’ are 27%27\%, 53%53\%, 65%65\% and 73%73\%, respectively, showing that fashion landmarks are the most effective representation for attribute prediction of fashion items. Fig.7 (a) shows the top-55 recall rates of ten representative attributes, e.g. ‘stripe’, ‘long-sleeve’, and ‘V-neck’. We observe that fashion landmark outperforms all the other localization schemes in all the attributes, especially for part-based attributes, such as ‘zip-up’ and ‘shoulder-straps’.

Conclusions

This paper introduced fashion landmark detection, which is an important step towards robust fashion recognition. To benchmark fashion landmark detection, we introduced a large-scale fashion landmark dataset (FLD). With FLD, we proposed a deep fashion alignment network (DFA) for robust fashion landmark detection, which leverages pseudo-labels and auto-routing mechanism to reduce the large variations presented in fashion images. Extensive experiments showed the effectiveness of different components as well as the generalization ability of DFA. To demonstrate the usefulness of fashion landmark, we evaluated on two fashion applications, clothing attribute prediction and clothes retrieval. Experiments revealed that fashion landmark is a more discriminative representation than clothes bounding boxes and human joints for fashion-related tasks, which we hope could facilitate future research.

Acknowledgements This work is partially supported by SenseTime Group Limited, the Hong Kong Innovation and Technology Support Programme, the General Research Fund sponsored by the Research Grants Council of the Kong Kong SAR (CUHK 416312), the External Cooperation Program of BIC, Chinese Academy of Sciences (No.172644KYSB20150019), the Science and Technology Planning Project of Guangdong Province (2015B010129013, 2014B050505017), and the National Natural Science Foundation of China (61503366, 61472410; Corresponding author is Ping Luo).

References

Appendix 0.A More Comparisons between DFA and DeepPose

More comparisons between DFA and DeepPose are given in Table 5.

For all the three stages, DFA employs the images as input while DeepPose uses the local image patches (i.e. the regions predicted by the previous stage) except stage-1. The last two stages of DFA also employ the estimated landmarks’ positions as additional input signal. As shown in the Fig.1 (b) and (c) in the paper, the local patches of fashion images have large variations of scale and appearance. They loss much of the discriminative information, such that using the local patches as input may harm the accuracy. For example, as demonstrated in Fig.9 (b) below, where the ‘right sleeve’ takes place in an unusual pose. Since the initial estimation in stage-1 is far from ground truth, it is hard for DeepPose (stage-2 and 3) to correct only based on the local observations.

DFA employs the full images as input other than local patches, without losing the context information. Furthermore, the estimated landmark positions of the previous stage and the pseudo labels implicitly capture global shape structure and local shape constraints respectively, which facilitate to improve performance. For example, as shown in Fig.9 (a), we see that pseudo-labels in stage-1 mainly reflect holistic pose like hold arms. Pseudo-labels in stage-2 instead suggest possible ways to correct initial estimations, e.g. stretching the estimations of ‘sleeve’ further to the end. Moreover, pseudo-labels in stage-3 incorporate pairwise relationships between fashion landmarks, e.g. the relative position from ‘left sleeve’ to ‘right sleeve’ here. Combining all three guidance leads to the effectiveness and robustness of DFA.

The carefully designed architecture of DFA reduces the number of VGGs in the network cascade, increasing efficiency of both the training and testing procedures.

Appendix 0.B Implementation Details

In this section, we introduce the implementation details of Deep Fashion Alignment (DFA) and DeepPose .

Both Deep Fashion Alignment (DFA) and DeepPose adopt the cascading framework. In each stage, VGG16 is used as the regression model li=Ψ(I;θ)l_{i}=\Psi\left(I;\theta\right), where II is the input image and θ\theta is the network parameters. We replace the original 10001000-way classification layer with one 1616-way landmark position regression layer and eight 33-way landmark visibility classification layers. Clothes bounding box is taken as the initial input while both landmark positions and visibility act as the supervision signals for training. Please note that the truncated landmarks doesn’t induce any error for the landmark position regression.

We also normalize the landmark positions following . Assume a clothes bounding box is defined by its center (xc,yc)\left(x_{c},y_{c}\right) as well as width bwb_{w} and height bhb_{h}: b=(xc,yc,bw,bh)b=\left(x_{c},y_{c},b_{w},b_{h}\right). Then our final landmark position li=(xi,yi)l_{i}=\left(x_{i},y_{i}\right) for landmark ii can be obtained by normalizing the absolute landmark coordinates liorig=(xiorig,yiorig)l_{i}^{orig}=\left(x_{i}^{orig},y_{i}^{orig}\right):

where N(⋅)\mathcal{N}\left(\cdot\right) is the normalization function. After inference, the estimated landmarks in absolute coordinates l^iorig\hat{l}_{i}^{orig} can be readily read as:

where N−1(⋅)\mathcal{N}^{-1}\left(\cdot\right) is the inverse operation of normalization function N(⋅)\mathcal{N}\left(\cdot\right). Next, we introduce the detailed implementation pipeline for DPA and DeepPose respectively.

B.2 Input Preparation for DeepPose

DeepPose (stage-1) takes the clothing bounding box as input and regresses all NN fashion landmarks within its receptive field. These estimated landmark positions are used for the input preparation in subsequent stages. Then, there will be NN models in DeepPose (stage-2). For each model in stage-2, the input is part bounding box cropped around the estimated fashion landmark (e.g. ‘left sleeve’). The task is to regress the offset for the underlying fashion landmark. The size of the bounding box is set to be 120×120120\times 120, such that more contextual information can be included. Similar procedure applies for stage-3 for further refinement.

Appendix 0.C More Results

Fig.10 and Fig.11 demonstrate more visual results on attribute prediction/clothes retrieval and fashion landmark detection, respectively. DFA is capable of handling complex variations under different scenarios.