Fully-adaptive Feature Sharing in Multi-Task Networks with Applications in Person Attribute Classification
Yongxi Lu, Abhishek Kumar, Shuangfei Zhai, Yu Cheng, Tara Javidi, Rogerio Feris
Introduction
Humans possess a natural yet remarkable ability of seamlessly transferring and sharing knowledge across multiple related domains while doing inference for a given task. Effective mechanisms for sharing relevant information across multiple prediction tasks (referred as multi-task learning) are also arguably crucial for making significant advances towards machine intelligence. In this paper, we propose a novel approach for multi-task learning in the context of deep neural networks for computer vision tasks. We particularly aim for two desirable characteristics in the proposed approach: (i) automatic learning of multi-task architectures based on branching, (ii) selective sharing among tasks with automated learning of whom to share with. In addition, we want our multi-task models to have low memory footprint and low latency during prediction (forward pass through the network).
A natural approach for enabling sharing across multiple tasks is to share model parameters (partially or fully) across the corresponding layers of the task-specific deep neural networks. At an extreme, we can imagine a fully shared multi-task network architecture where all layers are shared except the last layer which predicts the labels for individual tasks. However, this unrestricted sharing may suffer from the problem of negative transfer where inadequate sharing across two unrelated tasks can worsen the performance on both. To avoid this, most of the multi-task deep architectures share the bottom layers till some layer after which the sharing is blocked, resulting in task-specific sub-networks or branches beyond it . This is motivated by the observation made by several earlier works that bottom layers capture low level detailed features, which can be shared across multiple tasks, whereas top layers capture features at a higher level of abstraction that are more task specific. It can be further extended to a more general tree-like architecture, e.g., a smaller group of tasks can share parameters even after the first break-point at layer and breakup at a later layer. However, the space of such possible branching architectures is combinatorially large and current approaches largely make a decision based on limited manual exploration of this space, often biased by designer’s perception of the relationship among different tasks .
Our goal in this work is to develop a principled approach for designing multi-task deep learning architectures obviating the need for tedious manual explorations. The proposed approach operates in a greedy top-down manner, making branching and task-grouping decisions at each layer of the network using a novel criterion that promotes the creation of separate branches for unrelated tasks (or groups of tasks) while penalizing for the model complexity. Since we also desire a multi-task model with low memory footprint, the proposed approach starts with a thin network and dynamically grows it during the training phase by creating new branches based on the aforementioned criterion. We also propose a method based on simultaneous orthogonal matching pursuit (SOMP) for initializing a thin network from a pretrained wider network (e.g., VGG-16) as a side contribution in this work.
We evaluate the proposed approach on person attribute classification, where each attribute is considered a task (with non-mutually exclusive labels), achieving state-of-the-art results with highly compact multi-task models. On the CelebA dataset , we match the current top results on facial attribute classification (90% accuracy) with a model 90x more compact and 3x faster than the original VGG-16 model. We draw similar conclusions for clothing category recognition on the DeepFashion dataset , demonstrating that we can perform simultaneous facial and clothing attribute prediction using a single compact multi-task model, while preserving accuracy.
In summary, our main contributions are listed below:
We propose to automate learning of multi-task deep network architectures through a novel dynamic branching procedure, which makes task grouping decisions at each layer of the network (deciding with whom each task should share features) by taking into account both task relatedness and complexity of the model.
A novel method based on Simultaneous Orthogonal Matching Pursuit is proposed for initializing a thin network from a wider pre-trained network model, leading to faster convergence and higher accuracy.
We perform joint prediction of facial and clothing attributes, achieving state-of-the-art results on standard datasets with a significantly more compact and efficient multi-task model. We also conduct relevant ablation studies providing insights into the proposed approach.
Related Work
Multi-Task Learning. There is a long history of research in multi-task learning . Most proposed techniques assume that all tasks are related and appropriate for joint training. A few methods have addressed the problem of “with whom” each task should share features . These methods are generally designed for shallow classification models, while our work investigates feature sharing among tasks in hierarchical models such as deep neural networks.
Recently, several methods have been proposed for multi-task learning using deep neural networks. HyperFace simultaneously learns to perform face detection, landmarks localization, pose estimation and gender recognition. UberNet jointly learns low-, mid-, and high-level computer vision tasks using a compact network model. MultiNet exploits recurrent networks for transferring information across tasks. Cross-ResNet connects tasks through residual learning for knowledge transfer. However, all these methods rely on hand-designed network architectures composed of base layers that are shared across tasks and specialized branches that learn task-specific features.
As network architectures become deeper, defining the right level of feature sharing across tasks through handcrafted network branches is impractical. Cross-stitching networks have been recently proposed to learn an optimal combination of shared and task-specific representations. Although cross-stitching units connecting task-specific sub-networks are designed to learn the feature sharing among tasks, the size of the network grows linearly with the number of tasks, causing scalability issues. We instead propose a novel algorithm that makes decisions about branching based on task relatedness, while optimizing for the efficiency of the model. We note that other techniques such as HD-CNN and Network of Experts also group related classes to perform hierarchical classification, but these methods are not applicable for the multi-label setting (where labels are not mutually exclusive).
Model Compression and Acceleration. Existing deep convolutional neural network models are computationally and memory intensive, hindering their deployment in devices with low memory resources or in applications with strict latency requirements. Methods for compressing and accelerating convolutional networks include knowledge distillation , low-rank-factorization , pruning and quantization , structured matrices , and dynamic capacity networks . These methods are task-agnostic and therefore most of them are complementary to our approach, which seeks to obtain a compact multi-task model by widening a low-capacity network based on task relatedness. Moreover, many of these state-of-the-art compression techniques can be used to further reduce the size of our learned multi-task architectures.
Person Attribute Classification. Methods for recognizing attributes of people, such as facial and clothing attributes, have received increased attention in the past few years. In the visual surveillance domain, person attributes serve as features for improving person re-identification and enable search of suspects based on their description . In e-commerce applications, these attributes have proven effective in improving clothing retrieval , and fashion recommendation . It has also been shown that facial attribute prediction is helpful as an auxiliary task for improving face detection and face alignment .
State-of-the-art methods for person attribute prediction are based on deep convolutional neural networks . Most methods either train separate classifiers per attribute or perform joint learning with a fully shared network . Multi-task networks have been used with base layers that are shared across all attributes, and branches to encode task-specific features for each attribute category . However, in contrast to our work, the network branches are hand-designed and do not exploit the fact that some attributes are more related than others in order to determine the level of sharing among tasks in the network. Moreover, we show that our approach produces a single compact network that can predict both facial and clothing attributes simultaneously.
Methodology
Traditional approaches tackle the width design problem largely through hand-crafted layer design and manual model selection. Notably, popular deep convolutional network architectures, such as AlexNet , VGG , Inception and ResNet all use wider layers at the top of the network in what can be called an “inverse pyramid” pattern. These architectures serve as excellent reference designs in a myriad of domains, but researchers have noted that the width schedule (especially at the top layers) need to be tuned for the underlying set of tasks the network has to perform in order to achieve best accuracy .
Here we propose an algorithm that dynamically finds the appropriate width of the multi-task network along with the task groupings through a multi-round training procedure. It has three main phases:
Thin Model Initialization. We start with a thin neural network model, initializing it from a pre-trained wider VGG-16 model by selecting a subset of filters using simultaneous orthogonal matching pursuit (ref. Section 3.1).
Adaptive Model Widening. The thin initialized model goes through a multi-round widening and training procedure. The widening is done in a greedy top-down layer-wise manner starting from the top layer. For the current layer to be widened, our algorithm makes a decision on the number of branches to be created at this layer along with task assignments for each branch. The network architecture is frozen when the algorithm decides to create no further branches (ref. Section 3.2).
Training with the Final Model. In this last phase, the fixed final network is trained until convergence.
More technical details are discussed in the next few sections. Algorithm 1 provides a summary of the procedure.
The initial model we use is a thin version of the VGG-16 network. It has the same structure as VGG-16 except for the widths at each layer. We experiment with a range of thin models that are denoted as thin- models. The width of a convolutional layer of the thin- model is the minimum between and the width of the corresponding layer of the VGG-16 network. The width of the fully connected layers are set to . We shall call the “thinness factor”. Figure 1 illustrates a thin model side by side with VGG-16.
where is a truncated weight matrix that only keeps the rows indexed by the set . This problem is NP-hard, however, there exist approaches based on convex relaxation and greedy simultaneous orthogonal matching pursuit (SOMP) which can produce approximate solutions. We use the greedy SOMP to find the approximate solution which is then used to initialize the parameter matrix of the thin model as . We run this procedure layer by layer, starting from the input layer. At layer , after initializing , we replace with a column-truncated version that only keeps the columns indexed by to keep the input dimensions consistent. This initialization procedure is applicable for both convolutional and fully connected layers. See Algorithm 2.
2 Top-Down Layer-wise Model Widening
At the core of our training algorithm is a procedure that incrementally widens the current design in a layer-wise fashion. Let us introduce the concept of a “junction”. A junction is a point at which the network splits into two or more independent sub-networks. We shall call such a sub-network a “branch”. Each branch leads to a subset of prediction tasks performed by the full network. In the context of person attributes classification, each prediction is a sigmoid unit that produces a normalized confidence score on the existence of an attribute.
We propose to widen the network only at these junctions. More formally, consider a junction at layer with input and outputs . Note that each output is the input to one of the top sub-networks. Similar to Equation 1 the within-layer computation is given as
where parameterizes the connection from input to the ’th output at layer . The set is the indexing set . A junction is widened by creating new outputs at the layer below. To widen layer by a factor of , we make layer a junction with outputs. We use to denote an output in layer (each is an input for layer ) and to denote its parameter matrix. All of the newly-created parameter matrices have the same shape as (the parameter matrix before widening). The single output is replaced by a set of outputs where
Let be a given grouping function at layer . After widening, the within-layer computation at layer is given as (cf. Equation 3)
where the latter equality is a consequence of Equation 3. The widening operation sets the initial weight for to be equal to the original weight of . It allows the widened network to preserve the functional form of the smaller network, enabling faster training.
To put the widening of one junction into the context of the multi-round progressive model widening procedure, consider a situation where there are tasks. Before any widening, the output layer of the initial thin multi-task network has a junction with outputs, each is the output of a sub-network (branch). It is also the only junction at initialization. The widening operation naturally starts from the output layer (denoted as layer ). It will cluster the branches into groups where . In this manner the widening operation creates branches at layer . The operation is performed recursively in a top-down manner towards the lower layers. Note that each branch will be associated with a sub-set of tasks. There is a 1-1 correspondence between tasks and branches at the output layer, but the granularity goes coarser at lower layers. An illustration of this procedure can be found in Figure 2.
3 Task Grouping based on the Probability of Concurrently Simple or Difficult Examples
Ideally, dissimilar tasks are separated starting from a low layer, resulting in less sharing of features. For similar tasks the situation is the opposite. We observe that if an easy example for one task is typically a difficult example for another, intuitively a distinctive set of filters are required for each task to accurately model both in a single network. Thus we define the affinity between a pair of tasks as the probability of observing concurrently simple or difficult examples for the underlying pair of tasks from a random sample of the training data.
To make it mathematically concrete, we need to properly define the notion of a “difficult” and a “simple” example. Consider an arbitrary attribute classification task . Denote the prediction of the task for example as , and the error margin as , where is the binary label for task at sample . Following the previous discussion, it seems natural to set a fixed threshold on to decide whether example is simple or difficult. However, we observe that this is problematic since as the training progresses most of the examples will become simple as the error rate decreases, rendering this measure of affinity useless. An adaptive but universal (across all tasks) threshold is also problematic as it creates a bias that makes intrinsically easier tasks less related to all the other tasks.
The estimated task affinity is used directly for the clustering at the output layer. It is natural as branches at the output layer has a 1-1 map to the tasks. But at lower layers the mapping is one to many, as a branch can be associated with more than one tasks. In this case, affinity is computed to reflect groups of tasks. In particular, let , denote two branches at the current layer, where and denotes the -th and -th task associated with each branch respectively. The affinity of the two branches are defined by
4 Complexity-aware Width Selection
The number of branches to be created determines how much wider the network becomes after a widening operation. This number is determined by a loss function that balances complexity and the separation of dissimilar tasks to different branches. For each number of clusters , we perform spectral clustering to get a grouping function that associates the newly created branches with the old branches at one layer above. At layer the loss function is given by
where is a penalty term for creating branches at layer , is a penalty for separation. is defined as the number of pooling layers above the layer and is the unit cost for branch creation. The first term grows linearly with the number of branches, with a scalar that defines how expensive it is to create a branch at the current layer (which is heuristically set to double after every pooling layers). Note that in this formulation a larger encourages the creation of more branches. We call the branching factor. The network is widened by creating the number of branches that minimizes the loss function, or .
The separation term is a function of the branch affinity matrix . For each , we have
and the separation cost is the average across each newly created branches
Note Equation 10 measures the maximum distances (minimum affinity) between the tasks within the same group. It penalizes cases where very dissimilar tasks are included in the same branch.
Experiments
We perform an extensive evaluation of our approach on person attribute classification tasks. We use CelebA dataset for facial attribute classification tasks and Deepfashion for clothing category classification tasks. CelebA consists of images of celebrities labeled with 40 attribute classes. Most images also include the torso region in addition to the face. Our models are evaluated using the standard classification accuracy (average of classification accuracy rate over all attribute classes) and the top-10 recall rate (proportion of correctly retrieved attributes from the top-10 prediction scores for each image). Top-10 is used as there are on average about 9 positive facial attributes per image on this dataset. DeepFashion is richly labeled with 50 categories of clothes, such as “shorts”, “jeans”, “coats”, etc. (the labels are mutually exclusive). Faces are often visible on these images. We evaluate top-3 and top-5 classification accuracy to directly compare with benchmark results in .
We establish three baselines. The first baseline is a VGG-16 model initialized from the a model trained from imdb-wiki gender classification . The second baseline is a low-rank model with low rank factorization at all layers. This model is also initialized from the imdb-wiki gender pretrained model, but the initialization is through truncated Singular Value Decomposition (SVD) . The number of basis filters is 8-16-32-64-64 for the convolutional layers, 64-64 for the two fully-connected layers and 16 for the output layer. The third is a thin model initialized using the SOMP initialization method introduced in Section 3.1, using the same pre-trained model. Our VGG-16 baselines are stronger than all previously reported methods, while the low-rank baselines closely matches the state-of-the-art while being faster and more compact. The thin baseline is up to 6 times faster, 500 times more compact than the VGG-16 baseline, but still reasonably accurate.
We find several contributing factors to the strength of our baselines. Firstly, the choice of pre-trained model is critical. Most recent works use the VGG face descriptor, whereas in our work we use the pre-trained model from imdb-wiki . For the thin baseline, it is also important to use Batch Normalization (BN) . Without the adoption of BN layers the training error ceases to decrease after a small number of training iterations. We observe this phenomenon in both random initialization and SOMP initialization.
A comparison of the models generated by our adaptive widening algorithm with baseline results are shown in Table 1 and 2. Our “branching” models achieves similar or better accuracy compared to these state-of-the-art methods, while being much more compact and faster.
2 Cross-domain Training of Joint Person Attribute Network
To examine the ability of our approach in handling cross-domain tasks, we train a network that jointly predict facial and clothing attributes. The model is trained on the union of the two training sets. Note that the CelebA dataset is not annotated with clothing labels, and the Deepfashion dataset is not annotated with facial attribute labels. To augment the annotations for both datasets, we use the predictions provided by the baseline VGG-16 models as soft training targets. We demonstrate that the joint model is comparable to the state-of-the-art on both facial and clothing tasks, while being a much more efficient combined model rather than two separate models. The comparison between the joint models with the baselines is shown in Table 1 and 2.
3 Visual Validation of Task Grouping
We visually inspect the task groupings in the generated model. Figure 3 displays the actual task grouping in the Branch-32-2.0 model trained on CelebA. The grouping are often highly intuitive. For instance, “5-o-clock Shadow”, “Bushy Eyebrows” and “No Beard”, which all describe some forms of facial hairs, are grouped. The cluster with “Heavy Makeup”, “Pale Skin” and “Wearing Lipstick” is clearly related. Groupings at lower layers are also sensible. As an example, the group “Bags Under Eyes”, “Big Nose” and “Young” are joined by “Attractive” and “Receding Hairline” at fc6, probably because they all describe age cues. This is particularly interesting as no human intervention is involved in model generation.
4 Ablation Studies
What are the advantages of grouping similar tasks? We shuffle the correspondence between training targets and the output of the network for “Branch-32-2.0” model from CelebA and report the reduction in accuracies for each tasks. Both random and manual shuffling are tested but we only report the one from manual shuffling as they are similar. In particular, for manual shuffling we choose a new grouping of tasks so that the network separates many tasks that are originally in the same branch. Figure 4 summarizes our findings. Clearly grouping tasks according to similarity improves accuracy for most tasks.
Closer examination yields other interesting observations. The three tasks that actually benefit from the shuffling significantly (unlike most of the tasks), namely “wavy hair”, “wearing necklace” and “pointy nose” are all from the branch with the largest number of tasks. This is sensible as after the shuffling they are not forced to share filters with many other tasks. But other tasks from the same branch, namely “black hair” and “wearing earrings” are significantly improved from the original grouping. One possible explanation is that while grouping similar tasks allow them to benefit from multi-task learning, some tasks are intrinsically more difficult and require a wider branch. Our current design lacks the ability to change the width of a branch, which is an interesting future direction.
Sub-optimal use of pretrained network or smaller capacity? The gap in accuracy between Branch-32-2.0 and VGG-16 baseline can be caused by sub-optimal use of the pretrained model or the intrinsically smaller capacity of the former. To determine if both factors contribute to the gap, we compare training the Branch-32-2.0 model and VGG-16 from scratch on CelebA. As neither model benefit from the information from a pre-trained network, we expect a much smaller gap in accuracy if the sub-optimal use of the pretrained model is the main cause. Our results summarized in Table 3 suggest that the smaller capacity of the Branch-32-2.0 model is likely the main reason for the accuracy gap.
How does SOMP help the training? We compare training with and without this initialization using the Baseline-thin-32 model on CelebA, under identical training conditions. The evolution of training and validation accuracies are shown in Figure 5. Clearly, the network initialized with SOMP initialization converges faster and better than the one without SOMP initialization.
Conclusion
We have proposed a novel method for learning the structure of compact multi-task deep neural networks. Our method starts with a thin network model and expands it during training by means of a novel multi-round branching mechanism, which determines with whom each task shares features in each layer of the network, while penalizing for the complexity of the model. We demonstrated compelling results of the proposed approach on the problem of person attribute classification. As future work, we plan to adapt our approach to other related problems, such as incremental learning and domain adaptation.