Unsupervised Learning of Object Landmarks through Conditional Image Generation
Tomas Jakab, Ankush Gupta, Hakan Bilen, Andrea Vedaldi
Introduction
There is a growing interest in developing machine learning methods that have little or no dependence on manual supervision. In this paper, we consider in particular the problem of learning, without external annotations, detectors for the landmarks of object categories, such as the nose, the eyes, and the mouth of a face, or the hands, shoulders, and head of a human body.
Our approach learns landmarks by looking at images of deformable objects that differ by acquisition time and/or viewpoint. Such pairs may be extracted from video sequences or can be generated by randomly perturbing still images. Videos have been used before for self-supervision, often in the context of future frame prediction, where the goal is to generate future video frames by observing one or more past frames. A key difficulty in such approaches is the high degree of ambiguity that exists in predicting the motion of objects from past observations. In order to eliminate this ambiguity, we propose instead to condition generation on two images, a source (past) image and a target (future) image. The goal of the learned model is to reproduce the target image, given the source and target images as input. Clearly, without further constraints, this task is trivial. Thus, we pass the target through a tight bottleneck meant to distil the geometry of the object (fig. 1). We do so by constraining the resulting representation to encode spatial locations, as may be obtained by an object landmark detector. The source image and the encoded target image are then passed to a generator network which reconstructs the target. Minimising the reconstruction error encourages the model to learn landmark-like representations because landmarks can be used to encode the geometry of the object, which changes between source and target, while the appearance of the object, which is constant, can be obtained from the source image alone.
The key advantage of our method, compared to other works for unsupervised learning of landmarks, is the simplicity and generality of the formulation, which allows it to work well on data far more complex than previously used in unsupervised learning of object landmarks, \eglandmarks for the highly-articulated human body. In particular, unlike methods such as [45; 44; 55], we show that our method can learn from synthetically-generated image deformations as well as raw videos as it does not require access to information about correspondences, optical-flow, or transformation between images.
Furthermore, while image generation has been used extensively in unsupervised learning, especially in the context of (variational) auto-encoders and Generative Adversarial Networks (GANs ; see section 2), our approach has a key advantage over such methods. Namely, conditioning on both source and target images simplifies the generation task considerably, making it much easier to learn the generator network . The ensuing simplification means that we can adopt the direct approach of minimizing a perceptual loss as in , without resorting to more complex techniques like GANs. Empirically, we show that this still results in excellent image generation results and that, more importantly, semantically consistent landmark detectors are learned without manual supervision (section 4). Project code and details are available at: http://www.robots.ox.ac.uk/~vgg/research/unsupervised_landmarks/
Related work
The recent approaches of [45; 44] learn to extract landmarks based on the principles of equivariance and distinctiveness. In contrast to our work, these methods are not generative. Further, they rely on known correspondences between images obtained either through optical flow or synthetic transformations, and hence, cannot leverage video data directly. Since the principle of equivariance is orthogonal to our approach, it can be incorporated as an additional cue in our method.
Unsupervised learning of representations has traditionally been achieved using auto-encoders and restricted Boltzmann machines [14; 47; 15]. InfoGAN uses GANs to disentangle factors in the data by imposing a certain structure in the latent space. Our approach also works by imposing a latent structure, but using a conditional-encoder instead of an auto-encoder.
Learning representations using conditional image generation via a bottleneck was demonstrated by Xue \etal in variational auto-encoders, and by Whitney \etal using a discrete gating mechanism to combine representations of successive video frames. Denton \etal factor the pose and identity in videos through an adversarial loss on the pose embeddings. We instead design our bottleneck to explicitly shape the features to resemble the output of a landmark detector, without any adversarial training. Villegas \etal also generate future frames by extracting a representation of appearance and human pose, but, differently from us, require ground-truth pose annotations. Our method essentially inverts their analogy network to output landmarks given the source and target image pairs.
Several other generative methods [42; 40; 37; 48; 32] focus on video extrapolation. Srivastava \etal employ Long Short Term Memory (LSTM) networks to encode video sequences into fixed-length representation and decode it to reconstruct the input sequence. Vondrick \etal propose a GAN for videos, also with a spatio-temporal convolutional architecture that disentangles foreground and background to generate realistic frames. Video Pixel Networks estimate the discrete joint distribution of the pixel values in a video by encoding different modalities such as time, space and colour information. In contrast, we learn a structured embedding that explicitly encodes the spatial location of object landmarks.
A series of concurrent works propose similar methods for unsupervised learning of object structure. Shu \etal learn to factor a single object-category-specific image into an appearance template in a canonical coordinate system, and a deformation field which warps the template to reconstruct the input, as in an auto-encoder. They encourage this factorisation by controlling the size of the embeddings. Similarly, Wiles \etal learn a dense deformation field for faces but obtain the template from a second related image, as in our method. Suwajanakorn \etal learn 3D-keypoints for objects from two images which differ by a known 3D transformation, by enforcing equivariance . Finally, the method of Zhang \etal shares several similarities with ours, in that they also use image generation with the goal of learning landmarks. However, their method is based on generating a single image from itself using landmark-transported features. This, we show is insufficient to learn geometry and requires, as they do, to also incorporate the principle of equivariance . This is a key difference with our method, as ours results in a much simpler system that does not require to know the optical-flow/correspondences between images, and can learn from raw videos directly.
Method
We are interested in learning a function that captures the “structure” of the object in the image as a set of object landmarks. As a first approximation, assume that are coordinates , one per landmark.
In order to learn the map in an unsupervised manner, we consider the problem of conditional image generation. Namely, we wish to learn a generator function
such that the target image is reconstructed from the source image and the representation of the target image. In practice, we learn both functions and jointly to minimise the expected reconstruction loss Note that, if we do not restrict the form of , then a trivial solution to this problem is to learn identity mappings by setting and . However, given that has the “form” of a set of landmark detections, the model is strongly encouraged to learn those. This is explained next.
Third, each heatmap is replaced with a Gaussian-like function centred at with a small fixed standard deviation :
One may wonder whether this construction can be simplified by removing steps two and three and simply consider (possibly after re-normalisation) as the output of the encoder . The answer is that these steps, and especially eq. 1, ensure that very little information from is retained, which, as suggested above, is key to avoid degenerate solutions. Converting back to Gaussian landmarks in eq. 2, instead of just retaining 2D coordinates, ensures that the representation is still utilisable by the generator network.
In practice, we consider a separable variant of eq. 1 for computational efficiency. Namely, let be the two components of each pixel coordinate and write . Then we set
where and respectively. Figure 2 visualizes the source , target and generated images, as well as overlaid with the locations of the unsupervised landmarks . It also shows the heatmaps and marginalized separable softmax distributions on the top and left of each heatmap for keypoints.
2 Generator network using a perceptual loss
The goal of the generator network is to map the source image and the distilled version of the target image to a reconstruction of the latter. Thus the generator network is optimised to minimise a reconstruction error . The design of the reconstruction error is important for good performance. Nowadays the standard practice is to learn such a loss function using adversarial techniques, as exemplified in numerous variants of GANs. However, since the goal here is not generative modelling, but rather to induce a representation of the object geometry for reconstructing a specific target image (as in an auto-encoder), a simpler method may suffice.
Inspired by the excellent results for photo-realistic image synthesis of , we resort here to use the “content representation” or “perceptual” loss used successfully for various generative networks [12; 1; 9; 19; 27; 30; 31]. The perceptual loss compares a set of the activations extracted from multiple layers of a deep network for both the reference and the generated images, instead of the only raw pixel values. We define the loss as where is an off-the-shelf pre-trained neural network, for example VGG-19 , denotes the output of the -th sub-network (obtained by chopping at layer ). As our goal is to have a purely-unsupervised learning, we pre-train the network by using a self-supervised approach, namely colorising grayscale images .
Experiments
In section 4.1 we provide the details of the landmark detection and generator networks; a common architecture is used across all datasets. Next, we evaluate landmark detection accuracy on faces (section 4.2) and human-body (section 4.3). In section 4.4 we analyse the invariance of the learned landmarks to various nuisance factors, and finally in section 4.5 study the factorised representation of object style and geometry in the generator.
The landmark detector ingests the image to produce landmark heatmaps . It is composed of sequential blocks consisting of two convolutional layers each. All the layers use filters, except the first one which uses . Each block doubles the number of feature channels in the previous block, with 32 channels in the first one. The first layer in each block, except the first block, downsamples the input tensor using stride 2 convolution. The spatial size of the final output, outputting the heatmaps, is set to . Thus, due to downsampling, for a network with , blocks, the resolution of the input image is , resulting in tensor. A final convolutional layer maps this tensor to a tensor, with one layer per landmark. As described in section 3.1, these feature channels are then used to render 2D-Gaussian maps (with ).
Image generation network.
2 Learning facial landmarks
We explore extracting source-target image pairs using either (1) synthetic transformations, or (2) videos. In the first case, the pairs are obtained as by applying two random thin-plate-spline (TPS) [11; 49] warps to a given sample image . We use the 200k CelebA images after resizing them to resolution. The dataset provides annotations for 5 facial landmarks — eyes, nose and mouth corners, which we do not use for training. Following we exclude the images in MAFL test-set from the training split and generate synthetically-deformed pairs as in [45; 55], but the transformations themselves are not required for training. We discount the reconstruction loss in the regions of the warped image which lie outside the original image to avoid modelling irrelevant boundary artefacts.
In the second case, are two frames sampled from a video. We consider VoxCeleb , a large dataset of face tracks, consisting of 1251 celebrities speaking over 100k English language utterances. We use the standard training split and remove any overlapping identities which appear in the test sets of MAFL and AFLW. Pairs of frames from the same video, but possibly belonging to different utterances are randomly sampled for training. By using video data for training our models we eliminate the need for engineering synthetic data.
Qualitative results.
Figure 2 shows the learned heatmaps and source-target-reconstruction-keypoints quadruplets for synthetic transformations and videos. We note that the method extracts keypoints which consistently track facial features across deformation and identity changes (\eg, the green circle tracks the lower chin, and the light blue square lies between the eyes). The regressed semantic keypoints on the MAFL test set are visualised in fig. 3, where they are localised with high accuracy. Further, the target image is also reconstructed accurately.
Quantitative results.
We follow [45; 44] and use unsupervised keypoints learnt on CelebA and VoxCeleb to regress manually-annotated keypoints in the MAFL and AFLW test sets. We freeze the parameters of the unsupervised detector network () and learn a linear regressor (without bias) from our unsupervised keypoints to 5 manually-labelled ones from the respective training sets. Model selection is done using validation split of the training data.
We report results in terms of standard MSE normalised by the inter-ocular distance expressed as a percentage , and show a few regressed keypoints in fig. 3. Before evaluating on AFLW, we finetune our networks pre-trained on CelebA or VoxCeleb on the AFLW training set. We do not use any labels during finetuning.
Sample efficiency. Figure 3 reports the performance of detectors trained on CelebA as a function of the number of supervised examples used to translate from unsupervised to supervised keypoints. We note that is already sufficient for results comparable to the previous state-of-the-art (SoA) method of Thewlis \etal , and that performance almost saturates at (vs. 19,000 available training samples).
Vs. SoA. Table 1 compares our regression results to the SoA. We experiment regressing from unsupervised landmarks, using the self-supervised and the supervised perceptual loss networks; the number of samples used for regression is maxed out () to be consistent with previous works. On both MAFL and AFLW datasets, at and error respectively (for ), we significantly outperform all the supervised and unsupervised methods. Notably, we perform better than the concurrent work of Zhang \etal (MAFL: ; AFLW: ), while using a simpler method. When synthetic warps are removed from , so that the equivariance constraint cannot be employed, our method is significantly better ( vs on MAFL). We are also significantly better than many SoA supervised detectors [54; 41; 57] using only supervised training examples, which shows that the approach is very effective at exploiting the unlabelled data. Finally, training with VoxCeleb video frames degrades the performance due to domain gap; including a bias in the linear regressor improves the performance.
Ablation study.
In table 2 we present two ablation studies, first on the keypoint bottleneck, and second where we compare against adversarial and other image-reconstruction losses. For both the settings, we take the best performing model configuration for facial landmark detection on the MAFL dataset.
Keypoint bottleneck. The keypoint bottleneck has two functions: (1) it provides a differentiable and distributed representation of the location of landmarks, and (2) it restricts the information from the target image to spatial locations only. When the bottleneck is replaced with a generic low dimensional fully-connected layer (as in a conventional auto-encoder) the performance degrades significantly This is because the continuous vector embedding is not encouraged to encode geometry explicitly.
3 Learning human body landmarks
Articulated limbs make landmark localisation on human body significantly more challenging than faces. We consider two video datasets, BBC-Pose , and Human3.6M . BBC-Pose comprises of 20 one-hour long videos of sign-language signers with varied appearance, and dynamic background; the test set includes 1000 frames. The frames are annotated with 7 keypoints corresponding to head, wrists, elbows, and shoulders which, as for faces, we use only for quantitative evaluation, not for training. Human3.6M dataset contains videos of 11 actors in various poses, shot from multiple viewpoints. Image pairs are extracted by randomly sampling frames from the same video sequence, with the additional constraint of maintaining the time difference within the range 3-30 frames for Human3.6M. Loose crops around the subjects are extracted using the provided annotations and resized to pixels. Detectors for and keypoints are trained on Human3.6M and BBC-Pose respectively.
Qualitative results.
Figure 4 shows raw unsupervised keypoints and the regressed semantic ones on the BBC-Pose dataset. For each annotated keypoint, a maximally matching unsupervised keypoint is identified by solving bipartite linear assignment using mean distance as the cost. Regressed keypoints consistently track the annotated points. Figure 5 shows quadruplets, as for faces, as well as the discovered keypoints. All the keypoints lie on top of the human actors, and consistently track the body across identities and poses. However, the model cannot discern frontal and dorsal sides of the human body apart, possibly due to weak cues in the images, and no explicit constraints enforcing such consistency.
Quantitative results.
Figure 4 compares the accuracy of localising the 7 keypoints on BBC-Pose against supervised methods, for both self-supervised and supervised perceptual loss networks. The accuracy is computed as the the -age of points within a specified pixel distance . In this case, the top two supervised methods are better than our unsupervised approach, but we outperform [33; 53] using 1k training samples (vs. 10k); furthermore, methods such as are specialised for videos and leverage temporal smoothness. Training using the supervised perceptual loss is understandably better than using the self-supervised one. Performance is particularly good on parts such as the elbow.
4 Learning 3D object landmarks: pose, shape, and illumination invariance
We train our unsupervised keypoint detectors on the SmallNORB dataset, comprising 5 object categories with 10 object instances each, imaged from regularly spaced viewpoints and under different illumination conditions. We train category-specific detectors for keypoints using image-pairs from neighbouring viewpoints and show results in fig. 6 for car and airplane (see supplementary material for visualisation of other object categories). Keypoints most invariant to various factors are visualised. These landmarks are especially robust to changes in illumination and elevation angle. They are also invariant to smaller changes in azimuth (), but fail to generalise beyond that. Most interesting, they localise structurally similar regions, even when there is a large change in object shape (\eg fig. 6-(d)); such landmarks could thus be leveraged for viewpoint-invariant semantic matching.
5 Disentangling appearance and geometry
In fig. 7 we show that our method can be interpreted as disentangling appearance from geometry. Generator/ keypoint networks are trained on SVHN digits , AFLW faces, and Human3.6M people. The generator network is capable of retaining the geometry of an image, and substituting the style with any other image in the dataset, including unrelated image pairs never seen during training. For example, in the third column we re-render the number 3 by mixing its geometry with the appearance of the number 5. This generalises significantly from the training examples, which only consist of pairs of digits sampled from the same house number instance, sharing a common style.
Conclusions
In this paper we have shown that a simple network trained for conditional image generation can be utilised to induce, without manual supervision, a object landmark detectors. On faces, our method outperforms previous unsupervised as well as supervised methods for landmark detection. The method can also extend to much more challenging data, such as detecting landmarks of people, and diverse data, such as 3D objects and digits.
We are grateful for the support provided by EPSRC AIMS CDT, ERC 677195-IDIU, and the Clarendon Fund scholarship. We would like to thank James Thewlis for suggestions and support with code and data, and David Novotný and Triantafyllos Afouras for helpful advice.