A Global-Local Emebdding Module for Fashion Landmark Detection

Sumin Lee, Sungchan Oh, Chanho Jung, Changick Kim

Introduction

Visual fashion analysis has attracted research attention in recent years because of its huge potential usefulness in industry. With a development of large fashion datasets , deep learning-based methods have achieved significant progresses in many tasks, such as clothes classification , clothing retrieval , and clothes generation . The fashion analysis is a challenging task due to the large variation and non-rigid deformation of clothes in images. To deal with this issue, recent methods utilized fashion landmarks. As illustrated in Fig 1, the fashion landmarks are key-points describing clothing structures such as collars, sleeves, waistlines, and hemlines. Since these methods improved their performance considerably, fashion landmark detection is one of the key issues. For localizing fashion landmarks, Liu et al. utilized a regression method. Wang et al. designed a model to capture kinematic and symmetry grammar of clothing landmarks.

However, the fashion landmarks exhibit large spatial variances across poses, scales, and styles of clothing items. Due to this property, the models require to understand comprehensive semantic information of clothes for accurate landmark detection.

To address this issue, we propose a fashion landmark detection network with a global-local embedding module to exploit rich contextual knowledge for clothes. The global-local embedding module is specifically designed for embedding fashion landmark information by employing a non-local block followed by convolutions. Note that, in our proposed module, the non-local operation captures long-range global dependencies within a clothing image, and the following convolution operation enhances the local representation power of output features.

With this processing, the output features are well facilitated to have global as well as local contextual information for input clothing image. Lastly, by upsampling the feature map to the same size of the input fashion image, our network predicts high-resolution heatmaps for more accurate landmark localization. To demonstrate the strength of our proposed network, we conduct experiments by using two datasets, Deepfashion and FLD . Experimental results show that our network outperforms the state-of-the-art methods.

Methods

Our network consists of three parts: a feature extractor, a global-local embedding module, and an upsampling network. The overall architecture of the proposed method is illustrated in Fig. 2. First, an input fashion image is resized to 224×224224\times 224 and fed into the feature extractor. For the feature extractor, we use VGG-16 network except the last convolutional layer and initialize weights with parameters pretrained on ImageNet . After the conv4_3 layer of VGG-16, the global-local embedding module is employed to generate rich landmark embedding features. We describe details of the global-local embedding module in 2.2.

From generated rich features, the landmark localization network predicts landmark heatmaps. For more accurate landmark estimation, we produce high-resolution landmark heatmaps which have the same size to the input image. Each value of the heatmaps represents the probability that there is a landmark. To increase the spatial size of the feature map, we use several transposed convolutions with kernel size 4, stride 2, and padding 1. At the end of the upsampling network, we utilize a 1×11\times 1 convolution to produce a heatmap for each landmark.

2 Global-local embedding module

We introduce a global-local embedding module to exploit rich contextual knowledge of a clothing item. As shown in the bottom of Fig. 2, the global-local embedding module consists of a non-local block and two convolutional layers.

Let x\mathbf{x} denote input features of the global-local embedding module, which is the conv4_3 feature map in our network. The output feature is first represented by the non-local block. Compared to a conventional convolution operation, the non-local operation calculates long-range dependencies between any two different points, and sums up the weighted input features. This operation is formulated as:

where C(x)C(\mathbf{x}) is a normalization factor of x\mathbf{x}, and θ(⋅)\theta(\cdot), ϕ(⋅)\phi(\cdot), and g(⋅)g(\cdot) are 1×11\times 1 convolutions. We define the function ff as an embedded Gaussian function as follows:

By utilizing the non-local operation with a residual connection, we obtain the y^\mathbf{\hat{y}}:

where w(⋅)w(\cdot) is a 1×11\times 1 convolution. Finally, y^\mathbf{\hat{y}} has rich global knowledge of the clothing item by weighted with long-range relationship.

For locating landmark positions, it is necessary to assimilate the global-informative features in a local manner. To improve the local representation power of the output feature z\mathbf{z}, two convolutional layers perform a weighted sum in local neighborhoods:

where FiF_{i} denotes the ii-th convolutional layer including a 3×33\times 3 convolution, a batch normalization, and a nonlinear function (i.e. ReLU) sequentially. By this process, the output feature z\mathbf{z} is manipulated to have not only global but also local contextual information of the input clothing image.

We use sequentially kk global-local embedding modules to gradually produce more advanced deep feature representation. We empirically set kk as 2 for Deepfashion and 3 for FLD.

Experiments

We evaluate the proposed network on two datasets, DeepFashion and FLD . DeepFashion is a large fashion dataset with 289,222 images. Those images are composed of 209,222 images for training, 40,000 images for validation, and 40,000 images for testing. FLD is a fashion landmark dataset with 123,016 images with more diverse variations in poses, scales, and background. FLD images are divided into 83,033 images for training, 19,991 images for validation, and remaining 19,991 images for testing. Each image of both two datasets is annotated with bounding boxes and landmarks. There are 8 landmarks for full-body clothes, 6 landmarks for upper clothes, and 4 landmarks for lower clothes.

We adopt normalized error (NE) metric for evaluation of our proposed method and comparisons to the other state-of-the-art methods. NE is the L2\mathit{L}_{2} distance between predicted and ground-truth landmarks in the normalized coordinate space, formulated as:

where pip_{i} and p^i\hat{p}_{i} are ground-truth and predicted landmarks of ii-th sample respectively, NN is the number of samples, and hih_{i} and wiw_{i} are the height and width of ii-the sample.

2 Results

We conduct experiments on two large datasets and compare results of the proposed network with the state-of-the-art methods . Table 1 summaries comparison results. From the table, our proposed network outperforms all the competitors at 0.0568 on FLD and 0.0393 on DeepFashion. Among the landmarks, it is hard to discriminate between waistlines and hems landmarks which have the largest error rate in other methods. Our proposed method achieves performance improvements in detecting waistlines and hems landmarks. From these results, it is proved that our network learns more advanced feature representations for fashion landmark detection with the aid of the proposed global-local embedding module.

We also visualize landmark detection results in Fig. 3. We can see that the proposed model discriminates right- and left-side landmarks even in the back-view and side-view images. According to these results, it is demonstrated that our network is robust to view-point variations and deformations.

Conclusion

In this paper, we have proposed a fashion landmark detection network with a global-local embedding module. Our proposed global-local embedding module is based on a non-local operation and a convolution operation. Utilizing global-local embedding module facilitated to exploit not only global but also local contextual knowledge of a clothing item. We evaluated our method on two benchmark datasets, and achieved the state-of-the-art performance over recent methods. Experimental results demonstrated that our proposed method improves the feature representation of a clothing item for fashion landmark detection.

Acknowledgement

• This work was supported by Electronics and Telecommunications Research Institute(ETRI) grant funded by the Korean government. [19ZS1100, Core Technology Research for Self-Improving Artificial Intelligence System]

References