Unsupervised Discovery of Object Landmarks as Structural Representations

Yuting Zhang, Yijie Guo, Yixin Jin, Yijun Luo, Zhiyuan He, Honglak Lee

Introduction

Computer vision seeks to understand object structures that reflect the physical states of objects and show invariance to individual appearance changes. Such intrinsic structures can serve as intermediate representations for high-level visual understanding. However, manual annotations or designs of object structures (e.g., skeleton, semantic parts) are costly and barely available for most object categories, making the automatic representation learning of object structure an attractive solution to this challenge.

Modern neural networks can learn latent representations to effectively solve various vision problems, including image classification , segmentation , object detection , human pose estimation , 3D reconstruction , and image generation . Several existing studies observe that these representations naturally encode massive templates of particular visual patterns. However, little evidence suggests that deep neural networks can naturally conceptualize the intrinsic structures of an object category compactly and perceptibly.

We aim at learning the physical parameters of conceptualized object structures without supervision. As a typical representation of intrinsic structures, landmarks represent the spatial configuration of stable local semantics across different object instances of the same category. Thewlis et al. proposed an unsupervised method to locate landmarks at the places where a convolutional neural network can detect stable visual patterns with high spatial equivariance to image transformations. However, this method did not explicitly encourage the landmarks to appear at critical locations for image modeling.

This paper addresses the problem of discovering landmarks in a generic image modeling process. In particular, we take landmark discovery as an intermediate step for image autoencoding. To leverage the training signals from the landmark-based image decoder, gradients need to go through the landmark coordinates, which makes Thewlis et al. ’s non-differentiable formulation infeasible. With a different way to calculate landmark coordinates, the image decoding module can make the landmark configuration informative regarding image reconstruction. We also introduce additional regularization terms to enforce the desirable properties of the detected landmarks and to prevent the landmark coordinates from encoding irrelevant or redundant latent information.

Our contributions in this paper are as follows.

We develop a differentiable autoencoder framework for object landmark discovery, which allows the image decoder to propagate training signals back to the landmark detection module. We introduce several soft constraints to reflect the properties of landmarks, forcing the discovered representations to be valid landmarks.

The proposed method discovers visually meaningful landmarks without supervision for a variety of objects. It outperforms the state-of-the-art method regarding the accuracy of predicting manually-annotated landmarks using discovered landmarks, and it performs comparably to fully supervised landmark detectors trained with a significant amount of labeled data.

The discovered landmarks show strong discriminative performance in recognizing visual attributes.

Our landmark-based image decoder is useful for controllable image decoding, such as object shape manipulation and structure-conditioned image generation.

Related work

Parts are commonly used object structures in computer vision. The deformable part-based model learns object part configurations to optimize the object detection accuracy, where similar ideas are rooted in earlier constellation approaches . A recent method based on the deep neural network performs end-to-end learning of deformable mixture of parts for pose estimation. The recurrent architecture and spatial transformer network are also used to discover and refine object parts for fine-grained image classification . In addition, discriminative mid-level patches can be also discovered without explicit supervision . Object-part discovery based on subspace analysis and clustering techniques is also shown to improve neural-network-based image recognition . Unlike the approaches specific to discriminative tasks, our work focuses on learning landmarks for generic image modeling.

Learning structural representations.

To capture the intrinsic structures of objects, existing studies disentangle visual content into multiple factors of variations, like the camera viewpoint, motion, and identity. The physical parameters of these factors are, however, still embedded in non-perceptible latent representations. Methods based on multi-task learning can take conceptualized structures (e.g., landmarks, masks, depth) as additional outputs. These structures in this setting are designed by humans and require supervision to learn.

Learning explicit structures for image correspondence.

Object structures create correspondence among object instances. Colocalization realizes the coarsest level of object correspondence. In a finer granularity, AnchorNet learns object parts and their correspondence across different objects and categories. WarpNet corresponds images in the same class by estimating the parameter of a thin plate spline (TPS) transformation , and it can roughly reconstruct 3D point cloud using a single-view image. The 3D interpreter network utilizes 2D landmark annotations to discover 3D skeletons as the explicit structures of objects. Our discovered landmarks are denser than object parts and sparser than 3D points. These landmark representations are also more sensitive to precise locations and obtained without supervision.

Landmark discovery with equivariance.

Object structures like landmarks should be equivariant to image transformation, including object and camera motions. Using this property in 2D image domain, Rocco et al. proposed to discover TPS control points to match pairs of object images densely. Thewlis et al. tried to densely map different objects to a canonical coordinate that reflects object structures. Instead of learning dense correspondence, Thewlis et al. took the same equivariance property as the guidance to train deep neural networks for object landmark discovery without manual supervision. A similar idea was formulated differently using hand-crafted features in early work . In comparison, our method not only takes the equivariance as a constraint to ensure the validity of the landmarks, but also use a differentiable formulation to incorporate the landmark coordinates into a generic image modeling process. Moreover, our discovered landmarks are more predictive of manually annotated landmarks than those obtained by Thewlis et al. , and our method works on a broader range of object categories.

Image modeling with landmarks.

Many unsupervised deep learning techniques exist to model visual content, including stacked autoencoders (SAE) , variational autoencoders , generative adversarial networks (GAN) , and auto-regressive networks (e.g., PixelCNN ). The GAN- and PixelCNN-based image generators conditioned on given object landmarks are proposed in . In contrast, our method uses the SAE framework to automatically discover landmarks that are informative for unsupervised image modeling.

Landmark detection.

A vast amount of supervised landmark detection methods exist in the literature. For human faces, there are active appearance models , template-based methods , regression-based methods , and more recent methods based on deep neural networks . Landmark detection methods are also available for human bodies , and birds . We use our discovered landmarks to predict manually annotated landmarks and compare our method with some recent supervised models.

Autoencoding-based landmark discovery

We aim at automatically discovering landmarks as an explicit representation of visual content. We propose an autoencoder that encodes landmark coordinates as (a part of) the encoder outputs (Section 3.1). Without supervision from hand-crafted labels, we introduce several constraints to encourage the discovered landmark coordinates to reflect the visual concept that agrees with human perception (Section 3.2). The proposed constraints prevent landmark-based representations from degenerating to non-perceptible latent representations. Another pathway of the encoder extracts the local latent descriptor for each discovered landmark (Section 3.3). We use both the landmarks and the latent descriptors to reconstruct the input image (Section 3.4). This section presents the fully differentiable neural network architecture (Figure 1) and training objectives (Section 3.5) for landmark discovery and unsupervised image modeling.

We formulate landmark localization as the problem of detecting particular keypoints in the image . Specifically, each landmark has a corresponding detector, which convolutionally outputs a detection score map with the detected landmark located at the maximum. In this framework, we use a deep neural network to transform an image I\mathbf{I} to a (K+1)(K+1)-channel detection confidence map D∈W×H×(K+1)\mathbf{D}\in^{W\times H\times(K+1)}. This map detects KK landmarks, and the (K+1)(K+1)-th channel represents background. D\mathbf{D}’s resolution W×HW\times H can be either equal to or less than that of I\mathbf{I}, but they should have the same aspect ratio. Inspired by the success of the stacked hourglass network in human pose estimation , we propose a light-weighted hourglass-style network to get the raw detection score map

where the matrix Dk\mathbf{D}_{k} is the kk-th channel of D\mathbf{D}, and the scalar Dk(u,v)\mathbf{D}_{k}(u,v) is the value of Dk\mathbf{D}_{k} at the pixel (u,v)(u,v). Later, we also use the vector D(u,v)∈K+1\mathbf{D}(u,v)\in^{K+1} to denote the multi-channel values of D\mathbf{D} at (u,v)(u,v). The same notation convention applies to other tensors of three axes.

Taking Dk\mathbf{D}_{k} as a weighting map, we use the weighted mean coordinate as the location of the kk-th landmark, i.e.,

where ζk=∑v=1H∑u=1WDk(u,v)\zeta_{k}=\sum_{v=1}^{H}\sum_{u=1}^{W}\mathbf{D}_{k}(u,v) is the spatial normalization factor. This formulation enables back-propagating the gradient from the downstream neural network through the landmark coordinates unless Dk\mathbf{D}_{k}’s mass is totally concentrated in a single pixel or totally uniformly distributed, which rarely happens in practice. As a shorthand notation, we write the landmarks and landmark detector as

The left half of the blue pathway in Figure 1 illustrates the landmark detector.

2 Visual concept of landmarks

Separation constraint

Ideally, the autoencoder training objective can automatically encourage the KK landmarks to be distributed at different local regions so that the whole image can be reconstructed. However, the initial randomness can make the landmarks, defined as the mean coordinates weighted by D\mathbf{D} as in (3), all around the image center in the beginning of the training. This can lead to local optima from which the gradient descent may not escape (see Appendix F.2). To circumvent this difficulty, we introduce an explicit loss to spatially separate the landmarks:

Equivariance constraint

This loss function is well-defined when gg is known. Inspired by Thewlis et al. , we simulate gg by a thin plate spline (TPS) with random parameters. We use random translation, rotation, and scaling to determine the global affine component of the TPS; and, we spatially perturb a set of control points to determine the local TPS component. Besides the conventional way of selecting TPS control points at a predefined uniform grid (as used in ), we also take the landmarks detected by the current model as the control points to improve simulated transformation’s focus on key image patterns. The two sets of control points are alternatively used in each optimization iteration (see Appendix F.3 for details). Moreover, when training sample appear in the form of video, we can also take the dense motion flow as gg and the actual next frame as I′\mathbf{I}^{\prime}.

Cross-object correspondence

Our model does not explicitly ensure the semantic correspondence among the landmarks discovered on different object instances. The cross-object semantic stability of the landmarks mainly relies on the fact that visual patterns activating the same convolutional filter are likely to share semantic similarities.

3 Local latent descriptors

For simple images, like in MNIST (see results for MNIST in Appendix B), multiple landmarks can be enough to describe the object shapes. For most natural images, however, landmarks are insufficient to represent all visual content, so extra latent representations are needed to encode complementary information. Though necessary, the latent representations should not encode too much holistic information that can overwhelm the image structures reflected by the landmarks. Otherwise, the autoencoder would not provide enough driving force to localize landmarks at meaningful locations. To achieve this trade-off, we attach a low-dimensional local descriptor to each landmark.

An hourglass-style neural network (see Appendix G.2) is introduced to obtain a feature map F\mathbf{F}, which has the same size as the detection confidence map D\mathbf{D}:

Note that F\mathbf{F} is in a feature space shared among all landmarks and has SS channels.

For each landmark, we use an average pooling weighted by a soft mask centered at the landmark to extract the local feature in the shared space. In particular, we take D‾k\overline{\mathbf{D}}_{k}, which is the Gaussian approximation of the detection confidence map defined in (6), as the soft mask. Then, a learnable linear operator is introduced for each landmarks to map the feature representation into a lower-dimensional individual space. Thus, the latent descriptor for the kk-th landmark is

where C<SC<S. The landmark-specific linear operator enables each landmark descriptor to encode a particular pattern in limited bits. We can also use (10) to extract a low-dimensional background descriptor. Since it is unreasonable to approximate the background confidence map with a Gaussian distribution, we exactly set D‾K+1=DK+1/ζK+1\overline{\mathbf{D}}_{K+1}=\mathbf{D}_{K+1}/\zeta_{K+1}. Note that fk\mathbf{f}_{k} is differentiable regarding both the feature map and the detection confidence map.

4 Landmark-based decoder

Figure 1 (right half of the blue pathway) illustrates this.

where τ(⋅)\tau(\cdot) is the non-linear activation function. This is illustrated by the right half of the red pathway in Figure 1.

Let ⟦⋯⟧\llbracket\cdots\rrbracket be the channel-wise concatenation. We use another hourglass-style network to reconstruct the image

The gray pathway in Figure 1 illustrates the image decoder.

5 Overall training objective

Experiments

We evaluate our method on a variety of datasets, including CelebA and AFLW for human faces, the cat head dataset , a car dataset built from PASCAL 3D , shoe images from UT Zappos50k , human pose images from Human3.6M , MNIST (Appendix B), and animal images from AwA (Appendix D).

Section 4.1 describes the datasets and shows the qualitative results of landmark discovery. In Section 4.2, we use the discovered landmarks to predict human-annotated landmarks, and we take the landmark detection accuracy as an indicator of the quality of discovered landmark. Section 4.3 demonstrates that our discovered landmarks can serve as effective image representations to predict shape-related facial attributes on CelebA. In Section 4.3, we show that our decoding module and the automatically discovered landmarks can be used to manipulate the object shapes.

Following , we use all facial images in the CelebA training set excluding those also appearing in the MAFL the test setThe MAFL dataset is a subset of CelebA. (then 16,1962 images in total) to train models for landmark discovery. We use the MAFL testing set (1000 images) for all testing cases and reserve the MAFL training set (19,000 images) to train prediction models for manually-annotated landmarks. By default, we use the cropped and aligned images provided in the dataset.

As shown in Figure 2, our method can automatically discover facial landmarks at semantically meaningful and stable locations, such as the forehead center, eyes, eyebrows, nose, and mouth corners. Compared to Thewlis et al. ’s method, which results in a few significant errors, our method can locate landmarks more robustly against pose variations and occlusions. Interestingly, our method can work out-of-the-box on head-shoulder portraits without training on exactly the same type of images (Figure 3). Figure 4 shows that our method can also learn and detect a larger number (e.g., 30) of high-quality landmarks on unaligned facial images. Appendix E.1 shows more results.

AFLW

Face images in AFLW are cropped differently from CelebA. The landmark discovery models (both ours and Thewlis et al. ’s) are pretrained on CelebA and finetuned on the AFLW training set (10,122 images) for adaptation. Sampled results on the AFLW testing set (2,991 images) are in Appendix E.2.

Cat heads

Our model is trained on 7,747 cat head images and tested on 1,257 images. Compared to human faces, cat heads show more holistic appearance variations. As shown in Figure 5, our model can discover consistent landmarks (e.g., ears, nose, mouth) across different cat species and interestingly predict landmark locations under significant occlusion (the first image). Appendix E.3 shows more results.

Cars

We build the profile-view car dataset by cropping the car images from the PASCAL 3D dataset. This dataset has a limited number of samples (567 images for training and 63 images for testing). As shown in Figure 7, our method can still learn meaningful landmarks (e.g., the windshield, driver-side door, wheels, rear) using a relatively small training set. Note that we transform the 3D annotations of the cars to 2D landmarks, so this dataset is ready for quantitative evaluation. Appendix E.4 shows more results.

Shoes

We use the same setting as in (49,525, training images and 500 testing images). As shown in Figure 6, landmarks are detected at semantically stable locations for different types of shoes. Appendix E.5 shows more results.

Human3.6M

Human3.6M contains human activity videos in stable backgrounds. We use all 7 subjects in Human3.6M training set for our evaluation (6 for training and 1 for validation)Training subject IDs: S1,S5,S6,S7,S8,S9; Validation subject IDs: S11.. We consider six activities (direction, discussion, posing, waiting, greeting, walking), in which human bodies are in the upright direction most of the time, resulting in 796,648 image frames for training and 87,975 image frames for testing. We removed the background using the off-the-shelf unsupervised background subtraction method provided in the dataset. The human bodies are cropped and roughly aligned regarding the foot location so that the excessive background regions are removed.

Compared to previously mentioned object types, human bodies have much more shape variations. As shown in Figure 8, our method can discover roughly consistent landmarks across a range of poses. In particular, the landmarks at the head, back, waist, and legs are stable across images. The landmarks at the arms are relatively less consistent across different poses, but they are still at semantically meaningful locations. Since the human body appearances in the frontal and back views are similar, we do not expect our discovered landmarks to distinguish the left and right sides of the human body, which means that a landmark at the left leg in the frontal view can locate the right leg in the back view. Since the training data is in the video format, optical flows are used as a short-term self-supervision for the eqvuivariance constraint in (8). Appendix C describes more details and results for Human3.6M experiments.

2 Prediction of ground truth landmarks

Unsupervised landmark learning is useful because of its potential to discover object structures that are coherent with the human’s perception. We evaluate discovered landmarks’ quality by predicting manually-annotated landmarks. Specifically, we use a linear model without a bias term to regress from the discovered landmarks to the human-annotated landmarks. Ground truth landmark annotations are needed to train this linear regressor. Thewlis et al. extensively used random TPS to augment both discovered and labeled landmarks for training (on CelebA and ALFW). However, we do not use data augmentation for our method to minimize the complexity of training. Even in this case, our method shows stronger performance.

In Table 1a, we regress the landmarks discovered using the models trained on the CelebA training set to the 5 annotated landmarks. The landmark labels in either the CelebA training set or the much smaller MAFL training set are used to train the regressor. Our method is not sensitive to the decreased size of the labeled training set. It outperforms Thewlis et al. ’s by 55% decrease of the landmark detection error and Thewlis et al. ’s by 45%. Notably, we achieve this with 30 discovered landmarks while theirs uses 50 landmarks or dense object frames. Additionally, Table 2 demonstrates the consistent superiority of our method on the cat head dataset (7 target landmarks9 annotated landmarks in total. We do not use the 2 at the ears.), the car dataset (6 target landmarks), and Human3.6MSee Appendix C for details (32 target landmarks). Figure 9 illustrates the landmark regression results.

Competitive performance compared to fully supervised methods.

Putting the landmark discovery model together with the linear regressor, we obtain a detector of human-designed landmarks. Unlike fully supervised methods, our model is trainable with a huge amount of unlabeled data, and the linear regressor can be trained using a relatively small amount of labeled data within a few minutes. Table 1b demonstrates that our model outperforms previous unsupervised methods and off-the-shelf pretrained fully-supervised models on the MAFL and AFLW testing sets. On AFLW, we take the 5 always-visible landmarks as the regression target. All models reported are either trained on the MAFL training set or publicly available.

Landmark detection with few labeled samples.

Taking our model as a detector of manually annotated landmarks, we find that less than 200 samples are enough for our model to achieve less than 4% mean error on the MAFL testing set, which is better than the performance of TCDCN and MTCNN. Learning curves are provided in Appendix F.1.

Effectiveness of different loss terms.

Our method combines several loss terms in the training objective (15). Table 1c shows that the removal of any term can cause performance drop of our model. In particular, the removal of the separation loss can devastate the model, and more detailed discussion about this loss term is in Appendix F.2. Our new differentiable formulation of the landmark validity constraints can already lead to a lower landmark detection error than Thewlis et al. ’s. Adding the reconstruction loss can further improve the accuracy.

3 Visual attribute recognition

Landmarks reflect object shapes. We use our discovered landmarks as a feature representation to recognize the shape-related binary facial attributes (13 labeled attributes are found) on CelebA. We still take the MAFL testing set for the quantitative evaluation. A linear SVM is trained for each attribute on the CelebA training set. We also compare our landmark coordinates with pretrained FaceNet (InceptionV1) top-layer (128-dim) and top conv-layer (1792-dim) features for the attribute recognition task. As shown in Table 3, our discovered landmarks (60-dim) outperforms the FaceNet top-layer features for most attributes. The conv-layer features outperform our landmarks slightly but have a much higher dimension. Combining the landmark coordinates and the FaceNet features, higher accuracy is achieved. This suggests that the discovered landmarks are complementary to image features pretrained on classification tasks.

4 Image manipulation and generation

Our jointly trained image decoding module conditioned its outputs on the input landmarks and their latent descriptors. If the two conditions are disentangled, we should be able to manipulate the object shape without changing other appearance factors by adjusting only the landmarks; or, vice versa. Note that landmark-based image morphing is not a new topic, and landmark-based hierarchical image decoding has also been explored recently . However, these landmarks are all designed and annotated by humans. So far, little evidence has suggested that the automatically discovered landmarks are accurate and representative enough as a reliable condition for image generation.

In Figure 10, we synthesize flows to adjust the discovered landmarks of an input image. Fixing the landmark latent descriptors, we obtain realistic facial and human-body images whose shapes agree with the new landmarks. Other than the facial and body shape, then appearance factors of the input image are not visually changed. This result suggests that our image decoding module can synthesize realistic image using the landmarks learned without supervision, and it also suggests that our discovered landmarks have become an explicit representation disentangled from other factors of variations for image modeling. Implementation details and more results about unsupervised landmark-based face manipulation are available in Appendix A.

In Figure 11, instead of adjusting the landmark coordinates, we use the discovered landmarks of a reference image as the control signal to generate new facial images. Following the GAN framework , the latent representation of the generated image is randomly drawn from a prior distribution. As in Reed et al. , the landmark coordinates and latent representation are combined for image generation. We adopt BEGAN for the discriminator and training objective. In addition, we apply a cyclic loss for the landmark coordinates, which encourages the same landmarks to be detected on the generated images as on the reference image. Our results provide additional evidence on the usefulness of the discovered landmarks for image modeling. Implementation details are in Appendix G.5.

Conclusion

We address the problem of unsupervised object landmark discovery and take it as an intermediate step of image representation learning. In particular, a fully differentiable neural network architecture is proposed for determining the landmark coordinates, together with soft contraints to enforce the validity of the detected landmarks. The discovered landmarks are visually meaningful and quantitatively more relevant to human-designed landmarks. In our framework, the discovered landmarks are an explicit part of the learned image representations. They are disentangled from the latent representations of the other appearance factors. The landmark-based explicit representations not only provide an interface for manipulating the image generation process but also appear to be complementary to pretrained deep-neural-network features for solving discriminative tasks.

This work was supported in part by ONR N00014-13-1-0762, NSF CAREER IIS-1453651, and Sloan Research Fellowship.

References

Appendix A More details and results on face manipulation using unsupervised landmarks

The discovered landmarks constitute the explicitly structural part of the image representation learned by our model. They provide an interface for humans to manipulate the image representation intuitively. Our decoding module can generate realistic facial images using the landmark descriptors extracted from a given image and different sets of landmarks.

In addition to the results shown in the main paper, we provide more qualitative results for unsupervised landmark-based face manipulation in this section. We train models of 10, 20, and 30 landmarks. To evaluate our method on many target landmarks, we take the landmarks discovered from other images and the interpolation/extrapolation between the landmarks discovered on two images as the targets.

In this section, we show results for our 30-landmark model (results for our 10,20-landmark model are available as supplementary videos). Figure 12 and 13 show results for manipulating all 30 landmarks. Figure 14 and 15 show results for manipulating the 3 landmarks at the mouth. Figure 16 and 17 show results for manipulating the 5 landmarks at the mouth and jaw. Videos are available in the following folders for gradually morphing the landmarks from their original coordinates to the target by linear interpolation.

30-landmark models: videos/face.30landmark-model.manipulate-{all|mouth|mouthext}-landmarks

10,20-landmark models: videos/face.{10|20}landmark-model.manipulate-{all|mouth}-landmarks

Appendix B Our model without landmark descriptors on MNIST

We train our landmark discovery model without the landmark descriptor pathway on MNIST in two settings: one model for each digit (Appendix B.1) and one model for all ten digits (Appendix B.2). Using the model for all digits, we can perform geometrically meaningful morphing between different digits (Appendix B.3), e.g., morphing 2 to 9. Videos for the morphing process are available in the folder videos/mnist-morphing .

In this section, our landmark discovery model is trained for each digit independently. As shown in Figure 18, the discovered landmarks are consistent within each digit despite the shape variations.

B.2 Models for all digits

In this section, we train a single model for all ten digits together. In Figure 19, the corresponding landmarks are shown in the same color. The landmarks discovered on the same digits are consistent across different image. More interesting, the corresponding landmarks across different digits are also semantically consistent. For examples, the orange cross is always in the middle of a digits, the blue cross at the bottom, and the green one at the most left-top part of every digit.

B.3 Digit morphing using discovered landmarks

Our model trained on the mix of all ten digits can discover corresponding patterns among different digit categories. Using this model, we can perform geometrically meaningful cross-category morphing. Figure 20 and 21 illustrate the morphing process. Note that the landmark coordinates constitute the full image representation for MNIST digits. Videos for the morphing process are available in the folder videos/mnist-morphing .

Appendix C Details and more results on Human3.6M

On the Human3.6M dataset, we train our landmark discovery model on six actions: waiting, posing, greeting, directions, discussions, and walking. We report quantitative results on predicting the 32 annotated landmarksSome markers are close to each other (e.g., on each foot, there are two markers), so the effective locations annotated by the markers are less than 32 (around 16). (acquired by wearable markers) using our models trained for the mix of all six actions and each independent action. We also show qualitative results of the mixed-action model (trained for all six actions).

The Human3.6M training data is in the video format. We can calculate the optical flows between nearby frames and take them as self-supervision for the equivariance constraint defined in (8). Following the same notations, the two frames are I\mathbf{I} and I′\mathbf{I}^{\prime}, and the optical flows define the transformation g(⋅,⋅)g(\cdot,\cdot). In particular, we use the Farneback method in OpenCV to compute the dense optical flows at 5-frame intervals. We then accumulate two optical flow fields to calculate the optical flows at 10-frame intervals. The 10-frame-interval optical flow fields in addition to the random TPS transform are used in the equivariance constraint.

Using the above formulation, we can reduce the ground landmark prediction error from 4.914.91 (without optical flows) to 4.144.14 (with optical flows). For all experiments on Human3.6M, we use the optical flow as self-supervision for equivariance as described above.

C.2 Quantitative results

We compare our model with Thewlis et al. ’s unsupervised landmark discovery method regarding the annotated-landmark prediction accuracy. Both models discover 16 landmarks. The whole training set is used to train the linear mapping from the discovered landmarks to the annotated ones. As discussed in the main paper, we do not expect the two unsupervised methods to distinguish the frontal and back views. Thus, in the evaluation, we compute the errors against the original landmark annotations and its left-right-flipped counterpartFor examples, we swap the coordinates of the left-shoulder landmark and the right-shoulder landmark. , and then we choose the minimum value as the final error. Note that, when flipping the landmark annotations, the landmarks for the whole body are flipped simultaneously. As to the linear regressor training, we propose the following training strategy.

Figure out the rough orientations of the human body, heuristically. If more than 2/3 of the left-hand side annotated landmarks are to the right of the right-hand side annotated landmarks, the human is in the frontal view.

Train the regressor using the images with the landmark annotations in the frontal view. The other images are ignored in this step.

Use the aforementioned evaluation protocol to determine if the landmark annotations on other images should be flipped or not. The model is then retrained with all the training images.

Repeat the step 3 until the model is converged.

As shown in Table 4, our method outperforms Thewlis et al. ’s method significantly. We also report the results obtained by Newell et al. ’s supervised stacked hourglass network using their off-the-shelf pretrained 16-landmark model. Both unsupervised methods perform worse than the supervised stacked hourglass network. However, our model is unsupervised, and our neural network architectures are also smaller. We believe that our results show the potential of unsupervised methods for discovering complicated object structures.

C.3 Qualitative results

We train our model and Thewlis et al. ’s model on all six chosen actions and perform the annotated-landmark prediction. Figure 22 shows the side-by-side comparison. In general, our method visually outperforms Thewlis et al. ’s. Figure 23 shows landmark discovery examples, where our method outperforms Thewlis et al. ’s method very significantly.

Appendix D Results on animals of mixed species

On the animal-with-attributes (AwA) dataset , we choose the profile images from five animal categories (antelope, deer, moose, horse, zebra) and try to detect landmarks on these mixed species of animals. As shown in Figure 24, even though multiple species of animals with different appearance are mixed, our method can still find several consistent landmarks. For example, the yellow cross is always on the hoof, the orange cross always above the back and the light green cross at the buttock. The landmarks are consistently detected despite the significant variations in species, pose and individual appearance.

Appendix E More qualitative results on human faces, cat heads, cars, and shoes

In this section, we show more result of landmark discovery and ground truth landmark prediction compared with Thewlis et al. . All the shown images are randomly sampled from the test set.

E.2 AFLW

E.3 Cat heads

E.4 Cars

E.5 Shoes

Appendix F Ablative study

Taking our model as a detector of manually annotated landmarks, we find that less than 200 samples are enough for our model to achieve less than 4% mean error on the MAFL testing set, which is better than the performance of two popular off-the-shelf models. This result suggests that it is possible to train a high-accurate landmark detector using only a few labeled sample when sufficient unlabeled samples are given to train our unsupervised model. We show its performance versus the number of labeled samples in Figure 38.

F.2 Evolution of detection confidence map during training

Figure 39 shows the detection confidence maps of an input image at different training stage of our model. In the beginning, the heatmap shows random values over the whole image. As a result, the landmarks, defined as the mean coordinates weighted by the confidence maps, are all at the center of the image. As the training goes on, the values gradually becomes spatially concentrated. With the separation loss defined in (7), the peaked value of each channel of the confidence map can move to a different location. Without the separation loss, every channel can have a peaked value at the center of the images, resulting in degenerate landmarks.

F.3 TPS control points

For the random TPS in our equivariance constraint (defined in (8)), we both use the regular-grid control points and take the discovered landmarks in the current iteration as the control points. The two sets of control points are alternatively used in each optimization iteration with 7::3 chance. We do not exhaustively tune the ratio and keep it the same in all experiments. As shown in Figure 5, the performance of our model is fairly insensitive to this ratio when the other hyper-parameters are fixed. However, introducing the discovered landmarks as the TPS control points does benefit.

Note that the discovered landmarks are clustered at the center of the image (see discussion about the separation constraint in Section 3.2) and cannot serve as good TPS control points. As a result, we use only the regular-grid control points in the beginning and start to apply the previously mentioned ratio after training the model for several thousands of iterations.

Appendix G Implementation details

The main paper and this supplementary materials report results on several datasets. Table 6 summarizes the image size we used for each each dataset. In our landmark discovery formulation, we need to perform random TPS to calculate the equivariance constraint in (8). It requires the image to have large enough margins so that the foreground will not be out of image due to the random transformation. Table 6 also summarizes the image size after padding with the edge values.

For different datasets, we crop the foreground images and prepare input images as follows.

We started from the cropped-and-aligned images (218×\times178) in CelebA dataset, scaled them to 100×\times100 pixels and then cropped the 80×\times80 center patches.

The dataset provides annotated bounding boxes. We enlarged the bounding boxes on each image with a margin on the top (1/41/4 height of the original bounding box), margins on the left and right (1/81/8 width of the original bounding box) so that the cropped facial images look similar to the CelebA data. The cropped images were also scaled to 80×\times80.

The dataset provides ground truth landmarks on the cat head. We figured out the bounding box for cropping each image according to the landmarks and then scaled the cropped images to 80×\times80.

We scaled the shoe images (102×\times136) to 64×\times64, and padded the images to be 80×\times80 with white margins. We linearly transformed the color value range from $toto[0.1,0.9]$ to avoid a huge amount of saturated responses for the output layer with the sigmoid activation. The color was scaled back for visualization.

The PASCAL 3D dataset provides annotations for the orientation and landmarks of several objects. We cropped the profile car images according to the bounding box of the ground truth landmarks. Slight margins are added to the bounding box, and the cropped image is scaled to a square image without preserving the aspect ratio.

We manually annotated bounding boxes and orientation labels (i.e., left, right, frontal, back) for several types of animals. For each rectangular bounding box, we enlarged its shorter edge to make the box square, cropped the patch, and scaled it to 64×\times64. We ignored the images with frontal and back views, and we flipped the right-facing animals horizontally so that all animals in the image face to the left.

The human body was cropped using a square bounding box from the original video frames. We use the 3D landmarks provided in this dataset, acquired by wearable markers, as side information to roughly align the scales and foot locations of the human bodies in different images. The cropped square images are scaled to 128×\times128 pixels. We use the provided segmentation masks, obtained by an off-the-shelf unsupervised background removal method, to mask out the background image with gray color. For visualization, we show the gray color as white for the printing clarity.

G.2 Network architectures

In Section 3 of the main paper, we describe the key architectures of our model and leave some details unspecified. This section describes the detailed neural network architectures.

Figure 40 summarizes the hourglass-like architectures that we used for images of different sizes. The image padding is explaned and specified in Appendix G.1. In general, an hourglass architecture has mirrored encoding (high-resolution to low-resolution) and decoding (low-resolution and high-resolution) architectures. Skip-links, made up of convolutional layers, create shortcuts from the encoding feature maps and decoding feature maps of the same resolution. The responses of the skip-links are fused with the main stream decoding responses using element-wise addition. We use the max-pooling to reduce the feature map size for encoding, and we upsample feature maps by the nearest interpolation for decoding. For each linear layer, a batch normalization layer is followed, and LeakyReLU is the default activation function.

G.3 Training strategy

We use Adam with an initial learning rate of 0.0010.001 to optimize the neural network parameters. We set the training batch size to 16 or 32There is no significant difference in the performance when using either 16 or 32 as the training batch size. for 80×\times80 and 60×\times60 images and 8 for 128×\times128 images. The learning rate starts from 10−3{10}^{-3} and decreases to 10−4{10}^{-4} and 10−5{10}^{-5} later. For color images, we do random brightness and contrast jittering.

For batch normalization, the global mean and variance are computed using a random subset of the training set when the neural network training is done. Note that using the running average during training for the global mean and variance can hurt the performance.

To implement the equivariance constraint in (8), we use random TPS transformation to obtain warped input images (paired with the original images) in each training iteration. Taking the normalized image height and width as 11, the random transformation parameters are:

Global affine component. Uniform random translation in ±0.15\pm 0.15; Gaussian random rotation with the standard deviation of 10∘10^{\circ}; Gaussian random scaling in the base-22 logarithm scale with the standard deviation of 1.251.25.

Local TPS. Gaussian random translation with the standard deviation of 0.10.1 (regarding the regular-grid control points) or 0.050.05 (regarding the landmark control points).

G.4 Hyper-parameters

Table 7 summarizes the dataset-specific hyper-parameters. Note that our model is not sensitive to minor changes of the hyper-parameters, but adjusting the hyper-parameters for each dataset can improve the performance slightly.

G.5 Details about face generations using unsupervised landmarks

Figure 11 in the main paper shows results of generating facial images conditioned on our discovered landmarks. In this experiment, we fix our landmark discovery module and use it to detect landmarks on training images. For the decoding module, we take the detected landmarks as a given input condition and map an isotropic Gaussian random variable to the latent part of the image representation. Inspired by , we first use deconvolutional layers to get a feature map from the random variable and use convolutional layers to get another map of the same size from the reconstructed detection confidence map. We use element-wise multiplication and channel-wise concatenation to fuse the two into one feature map as the input of the decoding neural network. Thus, the way of calculating the input feature map of the decoding module is not the same as our landmark discovery model.

We use the boundary equilibrium GAN (BEGAN) framework to design the discriminator, which encourages the decoder to generate realistic images. More concretely, an autoencoder is trained as the energy function to distinguish the real and generated images. To make sure the generated images are consistent with the landmark condition, we first use our landmark discovery module to detect landmarks on them. We then take the L1 distance between them and those from the corresponding real images (i.e., input training images for getting the landmarks) as an extra training loss for the decoding module.

This experiment is mainly to show that our discovered landmark is accurate and meaningful enough for controllable image generation.