Deep Imbalanced Attribute Classification using Visual Attention Aggregation

Nikolaos Sarafianos, Xiang Xu, Ioannis A. Kakadiaris

Introduction

We set out to develop a method that, given an image of a human, predicts its visual attributes. We posed the following questions: (i) what are the challenges of this problem? (ii) what have other people done? and (iii) how should a simple yet effective solution to this problem look like? Human attributes are imbalanced in nature. Bald individuals with a mustache wearing glasses are 14 to 43 times less likely to appear in the CelebA dataset compared to people without these characteristics. Large-scale imbalanced datasets can lead to biased models, optimized to favor the majority classes while failing to identify the subtle discriminant features that are required to recognize the under-represented classes. Setting the class imbalance aside, an additional challenge is identifying which areas in the image provide class-discriminant information. Giving emphasis to the upper part of an image, where the face is located, for attributes such as “glasses” and to the bottom part for attributes such as “long pants” can increase the recognition performance as well as the interpretability of our models . This challenge is usually addressed using visual attention techniques that output saliency maps. However, in the human attribute estimation domain, attention ground-truth annotations are not available to learn such spatial attributions.

Learning from imbalanced data is a well-studied problem in machine learning and computer vision. Traditional solutions include over-sampling the minority classes or under-sampling the majority classes to compensate for the imbalanced class ratio and cost-sensitive learning where classification errors are penalized differently. Such approaches have been extensively used in the past but they suffer from some limitations. For example, over-sampling introduces redundant information making the models prone to over-fitting, whereas under-sampling may remove valuable discriminative information. Recent works with deep convolutional neural networks introduced a sampling procedure of triplets, quintuplets or clusters of samples that satisfy some properties in the feature-space and used them to regularize their models. However, sampling triplets is a computationally expensive procedure and the characteristics of the triplets in a batch-mode setup might vary significantly.

Modern visual attribute classification techniques rely either on contextual information , side information , curriculum learning strategies or visual attention mechanisms to accomplish their task. Although context and side information can increase the recognition accuracy, we believe that a simple solution should not rely on those. We argue that a solution to the deep imbalanced attribute classification problem should: (i) extract discriminative information, (ii) leverage visual information that is specific for each attribute, and (iii) handle class imbalance. Since, to the best of our knowledge, there is no method available with such characteristics, we developed an approach that uses (i) a pre-trained network for feature extraction, (ii) a weakly-supervised visual attention mechanism at multiple scales for attribute specific information, and (iii) a loss function that handles class imbalance and focuses on hard and uncertain samples. By simplifying the problem and addressing each one of its challenges, we were able to achieve state-of-the-art results in both WIDER-Attribute and PETA datasets, which are the most widely used in this domain.

In the deep learning era, most models are overly-complicated for what they aspire to achieve. Carefully developed, well established, accurate baselines are essential to measure our progress over time. Towards this direction, there have been a few works recently with well-performing yet simple baseline approaches in the fields of 3D human pose estimation , image classification , and person re-identification . Our main contribution is the design and analysis of an end-to-end neural-network architecture that can be easily reproduced, is easy to train and achieves state-of-the-art visual attribute classification results. This performance improvement originates from extracting and aggregating visual attention masks at different scales as well as establishing a loss function for imbalanced attributes as well as hard or uncertain samples. Through experiments, ablation studies and qualitative results we demonstrate that:

A simple visual attention mechanism with only attribute-level supervision (no ground-truth attention masks) can improve the classification performance by guiding the network to focus its resources to those spatial parts that contain information relevant to the input image.

Extracting visual attention masks from more than one stage of the network and aggregating the information at a score-level enables the model to learn more discriminant feature representations.

Accounting for class imbalance is essential during learning from large datasets. While assigning prior class weights can alleviate part of this problem, we observed that a weighted-variant of the focal loss works consistently better by handling imbalanced classes and at the same time focusing on hard examples.

Due to the lack of strong supervision, the attention masks result in attribute predictions with high variance across subsequent epochs. To prevent this from destabilizing training and degrading the performance we introduce an attention loss function, which penalizes predictions that originate from attention masks with high prediction variance.

Since this work aspires to serve as a bar in the visual attribute classification domain that future works may improve upon, we identify some sources of error that still prevail, and point out future research directions to address them that require further exploration.

Related Work

Visual Attributes: When we are interested in providing a description of an object or a human, we tend to rely on visual attributes to accomplish this task. From early works to more recent ones visual attributes have been studied extensively in computer vision. Due to its commercial applications and the abundance of available data, the clothing domain has received significant attention recently with methods ranging from transfer learning and domain adaptation to retrieval and forecasting . Some works rely on contextual information , or leverage side information (e.g., viewpoint) to improve the recognition performance . Others , assume the existence of a predefined connection between parts and attributes (e.g., hats are usually above the head and in the upper 20% of the image) which does not always hold true as depicted in Figure 1. Zhu et al. proposed to learn spatial regularizations using an attention mechanism on a final ResNet representation. Their attention module outputs an attention tensor per attribute which is then fed to a multi-label classification sub-network. However, none of the aforementioned approaches consider the class imbalance that exists in such datasets, which prevents them from accurately recognizing under-represented attributes such as wearing sunglasses.

Visual Attention: Visual attention can be interpreted as a mechanism of guiding the network to focus its resources on those spatial parts that contain information relevant to the input image. In computer vision applications, visual attribution is usually implemented as a gating function represented with a sigmoid activation or a spatial softmax and is placed on top of one or more convolutional layers with small kernels extracting high-level information. Several interesting works have appeared recently that demonstrate the efficiency of visual attention . For example, the harmonious attention of Li et al. consists of four subparts that extract hard-regional attention, soft-spatial, and channel attention to perform person re-identification. Deciding where to place the attention mechanism in the network is a topic of active research with several single-scale and multi-scale attention techniques in the literature. Das et al. , opted for a single attention module, whereas others extract saliency heatmaps at multiple-scales to build richer feature representations.

Deep Imbalanced Classification: Two works that address this problem in an attribute classification framework are the large margin local embedding (LMLE) method and the class rectification loss (CRL) . In LMLE, quintuplets were sampled that preserve locality across clusters and discrimination between classes and a new loss was introduced. Dong et al. demonstrated that a careful hard mining of triplets within the batch acts as an effective regularization which improves the recognition performance of imbalanced attributes. However, LMLE is prohibitively computationally expensive as it comprises an alternating scheme for cluster refinement and classification. In a follow-up work the authors address this limitation by replacing the quintuplets with clusters. CRL on the other hand, samples triplets within the batch, complicating the training process significantly, as the convergence and the performance heavily rely on the triplet selection. In addition, CRL adds a fully-connected layer for each attribute before the final classification layer, which increases the number of parameters that need to be learned. Both methods approach class imbalance purely as a machine learning problem without focusing on the visual traits of the images that correspond to these attributes. Class imbalance arises also in detection problems , where the foreground object (or face) covers a small part of the image. A simple yet very effective solution is focal loss , which uses a weighting scheme at an instance-level within the batch to penalize hard misclassified samples and assign near-zero weights to easily classified samples.

Methodology

Given an image of a human our goal is to predict its visual attributes. Specifically, our input consists of an image xx along with its corresponding labels y=[y1,y2,…,yC]Ty=[y^{1},y^{2},\dots,y^{C}]^{T} where CC is the total number of attributes and ycy^{c} a binary label that indicates the presence or absence of a particular attribute in the image. In this work, we experimented with both ResNets and DenseNets as backbone architectures and thus, we opted for the representations after the third and the fourth stage/block of layers. The concept of extracting attention information can be expanded to more spatial resolutions/scales besides two at the expense of learning additional parameters. We will thus refer to the first part of the networks (up to stage/block three) as ϕ1(⋅)\phi_{1}(\cdot) and to the part from there and until the classifier as ϕ2(⋅)\phi_{2}(\cdot). In our primary network, which unless otherwise specified is a ResNet-101 architecture (deep CNN module in Figure 2), given an image xx, we obtain three-dimensional feature representations:

For 224×224224\times 224 images the attention mechanism is placed on features of channel size FiF_{i} equal to 1,0241,024 and 2,0482,048 with spatial resolutions Hi×WiH_{i}\times W_{i} equal to 14×1414\times 14 and 7×77\times 7, respectively. Finally, the classifier of the primary network outputs logits y^p(x)=Wpk2(x)+bp\hat{y}_{p}(x)=W_{p}k_{2}(x)+b_{p} where (Wp,bp)(W_{p},b_{p}) are the parameters of the classification layer.

With simplicity in mind, our attention mechanism, depicted in Figure 3, consists of three stacked convolutional layers (along with batch-normalization and ReLU) with a kernel size equal to one. Due to the multi-label nature of the problem, the last convolutional layer maps the channels to the CC number of classes (i.e., attributes). This is different than most attention works (with one label per image) that extract saliency maps of the same spatial/channel size of the given feature representation. The attribute-specific attention maps zh,wcz_{h,w}^{c} are then spatially normalized to ah,wca_{h,w}^{c}using a spatial softmax operation:

where h,wh,w correspond to the height and width dimension and cc to the corresponding attribute label. The spatial softmax operation results in attention masks with the property ∑h,wah,wc=1\sum_{h,w}a_{h,w}^{c}=1 for each attribute cc and is used to force the model to focus its resources to the most relevant region of the image. We will refer to the attention mechanism comprising the three convolutional layers as A\mathcal{A} and thus, for each spatial resolution ii we first obtain unnormalized attentions Zi(x)=A(ki(x))Z_{i}(x)=\mathcal{A}(k_{i}(x)), which are then spatially normalized using Eq. (2) resulting in normalized attention masks Ai(x)A_{i}(x).

Following the work of Zhu et al. , we concurrently pass the feature representations to a single convolutional layer with CC channels (same as the number of classes) followed by a sigmoid function. The role of this branch is to assign weights to the attention maps based on label confidences and avoid learning from the attention masks when the label is absent. The weighted attention maps reflect both attribute information at different spatial locations and label confidences. We observed in our experiments that this confidence-weighting branch boosts the performance by a small amount and helps the attention mechanism learn better saliency heatmaps (Figure 3 right).

Combining the output saliency masks from different scales can be done either at a prediction level (i.e., averaging the logits) or at a feature level . However, aggregating the attention masks at a feature level provided consistently inferior performance. We believe that this is because the two attention mechanisms extract masks that give emphasis to different spatial regions which, when added together, fail to provide the classifier with attribute-discriminative information. Thus, we opted for the former approach and fed each confidence-weighted attention mask to a classifier to obtain logits y^ai\hat{y}_{a_{i}} of the attention module ii. The final attribute predictions of dimensionality 1×C1\times C for an image xx are then defined as y^=\nicefrac(y^p+y^a1+y^a2)3\hat{y}=\nicefrac{{(\hat{y}_{p}+\hat{y}_{a_{1}}+\hat{y}_{a_{2}})}}{{3}}.

2 Deep Imbalanced Classification

Using the output predictions of the primary model y^p\hat{y}_{p} which have the same dimensionality 1×C1\times C (i.e., one for each attribute), a straight-forward approach adopted by Zhu et al. is to train the whole network using the binary cross-entropy loss Lb\mathcal{L}_{b} as:

where (y^pc,yc)(\hat{y}_{p}^{c},y^{c}) correspond to the logit and ground-truth labels for attribute cc, and σ(⋅)\sigma(\cdot) is the sigmoid activation function. However, such a loss function ignores completely the class imbalance. Aiming to alleviate this problem both at a class- and at an instance-level, we propose to use for our primary model a weighted-variant of the focal loss defined as:

where γ\gamma is a hyper-parameter (set to 0.50.5), which controls the instance-level weighting based on the current prediction giving emphasis to the hard misclassified samples, and wc=e−acw_{c}=e^{-a_{c}}, where aca_{c} the prior class distribution of the cthc^{th} attribute as in .

Unlike the face attention networks , which learn the attention masks based on ground-truth facial bounding boxes, in the human attribute domain such information is not available. This means that the attention masks will be learned based on attribute-level supervisions yy. The attention masks of dimensionality Hi×Wi×FiH_{i}\times W_{i}\times F_{i} are fed to a classifier which outputs logits y^ai\hat{y}_{a_{i}} for each spatial resolution ii. To account for the weak supervision of the attention network, we decided to focus on the attention masks with high prediction variance. Similar to the work of Chang et al. , after some burn-in epochs in which Lb\mathcal{L}_{b} is used, we start collecting the history HH of the predictions pH(ys∣xs)p_{H}(y_{s}|x_{s}) for the sths^{th} sample and compute the standard deviation across time for each sample within the batch:

where tt corresponds to the current epoch, var^\widehat{var} to the prediction variance estimated in history Ht−1H^{t-1} and ∣Hst−1∣|H_{s}^{t-1}| the number of stored prediction probabilities. The loss for the attention-masks at level ii with attribute-level supervision for each sample ss is defined as:

Attention mask predictions with high standard deviation across time will be given higher weights in order to guide the network to learn those uncertain samples. Note that for memory reasons, our history comprises only the last five epochs and not the entire history of predictions. We believe that such a scheme makes intuitively more sense in a weakly-supervised application rather than the fully-supervised scenarios (such as MNIST or CIFAR) in the original paper . Finally, the total loss that is used to train our network end-to-end (the primary network and the two attention modules) is defined as:

where La1\mathcal{L}_{a_{1}} is applied to the first attention module that extracts saliency maps of spatial resolution 14×1414\times 14, and La2\mathcal{L}_{a_{2}} is similarly applied to the second attention module after the fourth stage of the primary network with spatial resolution of 7×77\times 7. Disentangling the two loss functions enables us to focus on different types of challenges separately. The weighted focal loss Lw\mathcal{L}_{w}, handles the prior class imbalance per attribute using the weight wcw_{c} and at the same time focuses on hard misclassified positive samples via the instance-level weights of the focal loss. The attention loss La\mathcal{L}_{a} penalizes predictions that originate from attention masks with high prediction variance.

Experiments

To assess our method, we performed experiments and ablation studies on the publicly available WIDER-Attribute and PETA datasets, which are the most widely used in this domain. The training details for both datasets are provided in the supplementary material.

Dataset Description and Evaluation Metrics: The WIDER-Attribute dataset contains 13,789 images with 57,524 bounding boxes of humans with 14 binary attribute annotations each. Besides “gender”, which is balanced, the rest of the attributes demonstrate class imbalance, which can reach 1:181:18 and 1:281:28 for attributes such as “face-mask” and “sunglasses”. Following the training protocol of , we used the human bounding box as an input to our model and mean average precision (mAP) results are reported.

Baselines: We evaluate our approach against all the methods that have been tested on the WIDER-Attribute dataset, namely R-CNN , R*CNN , DHC , CAM , VeSPA , SRN , and a fine-tuned ResNet-101 network . In addition, we transform the last part of the network to perform multi-task classification (MTL) by adding a fully-connected layer with 64 units for each attribute. This enables us to additionally evaluate against CRL by forming triplets within the batch using class-level hard samples. Note that DHC and R*CNN leverage additional contextual information (e.g., scene context or image parts) that intuitively should boost the performance and VeSPA, which jointly predicts the viewpoint along with the attributes, did not train its viewpoint prediction sub-network on the WIDER-Attribute dataset. In SRN , the validation set was included in the training (which results in 20%20\% more training data) and samples from the test set were used to obtain an idea about the training performance. In order to allow for a fair comparison with the rest of the methods, we re-implemented their method (which is why there is an asterisk next to their work in Table 1) and trained it only on the training set of the WIDER-Attribute dataset. The difference between the reported results and our re-implementation is 1.21.2 in terms of mAP which is reasonable given the access to approximately 20%20\% less training data.

Evaluation Results: Our proposed approach achieves state-of-the-art results on the WIDER dataset by improving upon the second best work by 1.3 in terms of mAP and by 2.7 over ResNet-101 which was our primary network. The larger improvements achieved by our algorithm are in imbalanced attributes such as “Sunglasses” or “Plaid” that have visual cues in the image which demonstrates the importance of handling class imbalance and using visual attention to identify important visual information in the image. DHC and R*CNN that use additional context information performed significantly worse but this is partially because they utilize smaller primary networks. Overall the proposed approach performs better than or equal than the rest of the literature in all but one attributes and comes second behind CAM at recognizing hats.

2 Ablation Studies on WIDER

In our first ablation study (Table 2 - left), we investigate to what extent the primary network affects the final performance. This is because it is commonplace that as architectures become deeper, the impact of individual add-on modules becomes less significant. We observe that (i) the difference between a ResNet-50 and a DesneNet-201 architecture is more than 2%2\% in terms of mAP, (ii) DenseNet-201, which is the highest performing primary network, is almost as good as SRN due to its effective feature aggregation and reuse, and (iii) the mAP of the proposed approach is 2.12.1 more than the best performing primary network. In our second ablation study (Table 2 - right), we assess how each proposed component of our approach contributes to the final mAP. Our ResNet-101 baseline (w/o any class weighting) achieves 83.7% mAP which increases to 84.0% when the class weights are added. When the instance-level weighting is added (i.e., LwL_{w}) the total performance increases to 84.4%. These results indicate that it is important to take both class-level and instance-level weighting into consideration during imbalanced learning. Handling class imbalance using the weighted focal loss and adding our attention mechanism just at a single scale result in mAP equal to 85.085.0 which performs almost as well as the existing state-of-the-art. Adding the attention loss that penalizes attention masks with high prediction variance and expanding our attention module to two scales improves the final mAP to 86.4.

Qualitative Results: Figure 4, depicts attention masks for six successful (left) and three failure cases (right). We observe that for imbalanced attributes such as sunglasses that have discriminant visual cues, our attention mechanism locates successfully the corresponding regions, which explains the 7%7\% relative improved mAP for this attribute compared to our primary ResNet architecture.

3 Results on PETA

Dataset Description and Evaluation Metrics: The PETA dataset is a collection of 10 person surveillance datasets and consists of 19,000 cropped images along with 61 binary and 5 multi-value attributes. We used the same train/validation/test splits with the method of Sarfraz et al. and followed the established protocol of this dataset by reporting results on the 35 attributes for which the ratio of positive labels is higher than 5%. For the PETA dataset, two different types of metrics are reported namely label-based and example-based . For the label-based metrics due to the imbalanced class distribution, we used the balanced mean accuracy (mA) for each attribute that computes separately the classification accuracy of the positive and the negative examples and then computes the average. For the label-based metrics, we report accuracy, precision, recall, and F1-score averaged across all examples in the test set.

Baselines: We compared our approach with all the methods that have been tested on the PETA dataset, namely the ACN , DeepMAR , two variations of WPAL , VeSPA , the GoogleNet baseline reported by Sarfraz et al. , ResNet-101 and SRN .

Evaluation Results: From the complete evaluation results in Table 3, we observe that the proposed approach achieves state-of-the-art results in all example-based metrics and comes second to WPAL in terms of balanced mean accuracy (mA). We believe this is due to the fact that different methods use different metrics, based on which they optimize their models. For example, our approach is optimized based on the F1 score which balances between precision and recall and is applicable in search applications. Our approach improves upon a fine-tuned ResNet-101 architecture by approximately 2%2\% in terms of F1 score which demonstrates the importance of the visual attention mechanisms. Notably, we improve upon VeSPA in all evaluation metrics despite the fact that they utilize additional viewpoint information to train their model. Finally, we observe that by using the weighted variant of focal loss (Lw\mathcal{L}_{w}) instead of the binary-cross entropy loss (Lb\mathcal{L}_{b}), the F1 score of SRN increases by 1.7%1.7\%. This demonstrates why failing to account for class imbalance affects the performance of deep attribute classification models.

4 Ablation Studies on PETA

Based on our analysis an important question arises: can we achieve similar results with significantly fewer parameters? Aiming to find out the impact of large backbone architectures in the final performance, we investigated how each component of our work performs using a pre-trained DenseNet-121 architecture. DenseNet-121 contains 7.5×7.5\times less parameters compared to ResNet-101 due to efficient feature propagation and reuse. To our surprise, when all components are included (last row in Table 4), the performance drop in terms of F1 score is less than 2%2\%. In addition, we explored a variety of feature aggregations by either up-sampling the smaller attention masks, max-pooling the larger or mapping the larger to the smaller using a convolutional layer with a stride equal to two. Although the latter approach performed better than up-sampling/down-sampling, we observed that the aggregation of the attention information at a logit level is superior compared to feature level aggregation. We believe that this is because the two attention mechanisms extract masks that give emphasis to different spatial regions that when added together fail to provide the classifier with attribute-discriminative information.

5 Sources of Error and Further Improvements

Where does the proposed method fail and what are the characteristics of the failure cases? Aiming to gain a better understanding we will discuss separately the errors originating from the noise inherent to the input data and the errors related to modeling. A significant limitation of most pedestrian attribute classification methods (including ours) is that they resize the input data to a fixed square-size resolution (e.g., 224×224224\times 224) in order to feed them to deep pre-trained architectures. Human crops are usually rectangular captured from different viewpoints and thus, when they are resized to a square, important spatial information is lost. One possible solution to this would be feeding the whole image (before performing the human crop) at a fixed resolution that does not destroy the spatial relations and then extract human-related features using ROI-pooling at a stage within the network. To cope with the high viewpoint variance, the spatial transformer networks of Jaderberg et al. could be employed to align the input image before feeding it to the network, a practice which is very common in face recognition applications . A second source of error is the very low resolution of several images especially in the PETA dataset, which makes it hard even for the human eye to identify the attribute traits of the depicted human. Some training examples that demonstrate these sources of error are depicted in Figure 5. In addition, the provided annotations contain a third unspecified/uncertain class, which is used as negative during training in the literature, that further dilutes the learning process. Applying modern super-resolution techniques could alleviate this issue but only to some extent. Regarding errors due to modeling richer feature representations could be extracted using feature pyramid networks since they extract high-level semantic feature maps at multiple scales. Because the goal of this paper was to introduce a simple yet effective attribute classification solution, we refrained from building a complicated attention mechanism with a high number of parameters. Modern visual attention mechanisms could be adapted to a multi-label setup and applied to achieve superior performance at the expense of a larger parameter space.

Conclusion

Learning the visual attributes of humans is a multi-label classification problem that suffers from large class imbalance and lack of semantic/spatial attribute annotations. To address these challenges, we developed a simple yet effective and easy-to-reproduce architecture that outputs visual attention masks at multiple scales and handles effectively class imbalance and samples with high prediction variance. We introduced a weighted variant of focal loss that handles the prior class imbalance per attribute and focuses on hard misclassified positive samples. In addition, we observed that the weakly-supervised attention masks result in high prediction variance and thus, we introduced an attention loss that penalizes accordingly such predictions. By simplifying the problem and addressing each one of its challenges, we achieve state-of-the-art results in both the WIDER-Attribute and PETA datasets, which are the most widely used in this domain. This work aspires to serve as a bar in the visual attribute classification domain that future works can improve upon. To facilitate this process, we performed ablation studies, identified some sources of error that still exist and pointed out possible future research directions that require further exploration.

Acknowledgments: This work has been funded in part by the UH Hugh Roy and Lillie Cranz Cullen Endowment Fund. All statements of fact, opinion or conclusions contained herein are those of the authors and should not be construed as representing the official views or policies of the sponsors.

References

Supplementary Material

Since in both datasets we used a pre-trained primary network we first froze its weights and learned the attention masks using their corresponding loss function. This was done, to avoid back-propagating large prediction errors from the attention masks to the pre-trained network. After a few epochs of training solely the attention mechanism, the primary network is then unfrozen and trained end-to-end to produce multi-attribute predictions. For the WIDER-Attribute dataset we set the learning rate equal to 0.0010.001 and use SGD with momentum set to 0.90.9 and a weight decay equal to 0.00050.0005. The learning rate was divided by 1010 (until 0.000010.00001) when the error plateaus in the validation set. During pre-processing, we resized all images to 256×256256\times 256 and extracted random crops of $(alongwithrandommirroringanddatashuffling)whichwerethenresizedto(along with random mirroring and data shuffling) which were then resized to224\times 224andprovidedasaninputtothenetwork.ForthePETAdatasetweusedAdamsinceitconsistentlyoutperformedSGDwithastartinglearningrateequaltoand provided as an input to the network. For the PETA dataset we used Adam since it consistently outperformed SGD with a starting learning rate equal to0.0001withthesameweightdecaybutwithlargercrops(intherangewith the same weight decay but with larger crops (in the range).Inbothdatasets,thebatchsizewassetto). In both datasets, the batch size was set to32$. We used MXNet/Gluon as our deep learning framework and a single NVIDIA GeForce GTX 1080 Ti GPU.

Architecture Details

Our backbone architecture is a ResNet-101 that extracts feature representations of dimensionality 7×7×20487\times 7\times 2048 which are then fed to a fully-connected layer. Its dimensionality is equal to the number of classes denoted by Cl\mathcal{C}_{l} which for the WIDER dataset is equal to 14. The attention modules are placed on “stage3_activation22” and “stage4_activation2”. Let Ck denote a Convolution-BatchNorm-ReLU layer with k filters and kernel size equal to 1 and Dk a fully-connected layer with k neurons. The attention module consists of C256-C256 and a convolutional layer with Cl\mathcal{C}_{l} number of filters. Its output is first spatially normalized and then multiplied by the output of the confidence weighting layer which is simply a convolutional layer with Cl\mathcal{C}_{l} number of filters and a sigmoid activation function. The output of the attention modules is fed to a C256-C512-C512-DCl\mathcal{C}_{l} subnetwork the last convolutional layer of which has a kernel size equal to the spatial dimensions. All layers are initialized with Xavier initialization.