Seesaw Loss for Long-Tailed Instance Segmentation
Jiaqi Wang, Wenwei Zhang, Yuhang Zang, Yuhang Cao, Jiangmiao Pang, Tao Gong, Kai Chen, Ziwei Liu, Chen Change Loy, Dahua Lin
Introduction
Deep learning-based object detection and instance segmentation approaches have achieved immense success on datasets with relatively balanced category distribution, e.g., COCO dataset . However, the distribution of categories in the real world is long-tailed . There are a few head classes containing abundant instances, while most other classes comprise relatively few instances.
On long-tailed datasets, existing instance segmentation frameworks fail to perform as accurately as on the datasets with balanced category distribution, exhibiting unsatisfactory performance on tail classes. Figure 1 shows the classification accuracy and instance segmentation performance of Mask R-CNN on LVIS dataset. The classifier in Mask R-CNN trained by Cross-Entropy Loss tends to misclassify tail categories as backgrounds or other confusing head classes, which leads to extremely low accuracy on tail classes.
The primary reason for this undesired phenomenon is that the instances from head classes are predominant in a long-tailed dataset. These instances contribute an overwhelmingly large quantity of negative samples for tail classes. Thus, the gradients of positive and negative samples on a tail class are heavily imbalanced, leading to a biased learning process for the classifier. One can imagine that gradients of positive and negative samples resemble two objects positioned on each end of a seesaw (see Fig. 1). To balance them, a viable solution is to shorten the arm of the heavier end in the seesaw, which is equivalent to scaling down the overwhelming gradients of negative samples on the tail class by a factor. Nevertheless, blindly reducing the gradients of negative samples increases the risk of inducing false positives of tail classes, since samples of other classes are less punished when they are misclassified as tail classes. Thus, a specialized mechanism is needed to compensate for the excessively reduced penalties on tail classes.
In this work, we propose Seesaw Loss that dynamically re-balances positive and negative gradients for each category with two complementary factors, i.e., mitigation factor and compensation factor. According to the ratio between categories’ cumulative sample numbers during training, the mitigation factor reduces the penalty to relatively rare classes. When a false positive sample of one category is observed, the compensation factor will increase the penalty to that category. The synergy of the two above factors enables Seesaw Loss to mitigate the overwhelming punishments to tail classes as well as compensate for the risk of misclassification caused by diminished penalties.
Seesaw Loss has three appealing properties. 1) Seesaw Loss is dynamic. It explores the ratios of cumulative training sample numbers between different categories and instance-wise misclassification during training. This differs significantly to previous solutions that rely either on static group split or loss reweighting with constant values . 2) Seesaw Loss is self-calibrated. The mitigation and the compensation factor synergize to relieve the overwhelming punishments on tail classes as well as avoid increasing false positives of tail categories. On the contrary, previous methods blindly reduce punishments on tail classes or decrease the loss weights of head categories . 3) Seesaw Loss is distribution-agnostic. It does not rely on pre-computed datasets’ distribution , and it can operate well with any data sampler . By accumulating the number of samples in each class, Seesaw Loss gradually approximates the real data distribution during training to achieve more accurate balancing.
Through extensive experiments, we show consistent improvements of Seesaw Loss in different instance segmentation frameworks and data samplers. On the challenging LVIS dataset, Seesaw Loss achieves significant improvements of 6.0% AP and 2.1% AP upon Mask R-CNN with random sampler and repeat factor sampler , respectively. Even if switching to the stronger Cascade Mask R-CNN , we still observe an impressive improvement of 6.4% AP and 2.3% AP with random sampler and repeat factor sampler. To show the versatility of Seesaw Loss, we integrate it into the long-tailed image classification task. Seesaw Loss significantly improves the classification accuracy by 6% on ImageNet-LT dataset. Besides, we also explore the necessity of the decoupling training pipeline in Seesaw Loss. Experimental results demonstrate that Seesaw Loss provides a simpler and more effective solution to long-tailed instance segmentation without relying on complex training pipelines.
Related Work
Object Detection. Recent years have witnessed a remarkable improvement in object detection . A leading paradigm in this area is the two-stage pipeline , where the first stage generates a set of region proposals, and then the second stage classifies and refines the proposals. Unlike the two-stage approaches, the single-stage pipeline directly predicts bounding boxes. Classical single-stage approaches require densely populated anchors as a prior, while anchor-free methods manage to achieve similar or better performance without such prior. There are also attempts to apply cascade architecture to refine the bounding boxes’ predictions progressively.
Instance Segmentation. Instance segmentation is becoming popular in tandem with a surge in the interest in object detection. Early methods perform segmentation before object recognition . Via adding a mask prediction branch in the Faster R-CNN architecture, Mask R-CNN bridges the gap between object detection and instance segmentation. The idea is also adopted by in their cascading frameworks. More recent works introduce an even shorter pipeline by skipping the detection process and directly predicting mask for each instance. Seesaw Loss can easily cooperates with object detection and instance segmentation frameworks for the long-tailed datasets.
Long-Tailed Recognition. Long-tailed recognition tasks receive growing attention recently as the problems are closer to real-world applications. One representative solution to the problem is loss re-weighting . Loss re-weighting methods adopt different re-weighting strategies to adjust the loss of different classes based on each class’s statistics . Other common approches re-balance the distribution of the instance numbers in each class, e.g., repeat factor sampling and class-balanced sampling , both are based on the sample numbers of classes. Different sampling strategies can be adopted at different training stages to formulate a multi-stage training procedure . A recent work proposes a decoupling training pipeline. which first trains a good representation network with natural sampling and then finetunes the classifier with class-balanced sampling. There are also attempts to modify the classifier to improve the performance on tail classes, e.g., using different classifiers for different groups of classes , or use two classifiers trained with different data samplers .
Methodology
The classifier trained by the widely applied Cross-Entropy (CE) Loss (Sec. 3.1) is highly biased on long-tailed datasets, resulting in much lower accuracy of tail classes than head classes. The major reason is that gradients brought by positive samples are overwhelmed by gradients from negative samples on tail classes. Therefore, we propose Seesaw Loss to mitigate the overwhelming gradients of negative samples on tail classes as well as compensate the gradients of misclassified samples to avoid false positives (Sec. 3.2). We also explore some practical component designs to adopt Seesaw Loss in instance segmentation (Sec. 3.3).
We first revisit the most widely adopted Cross-Entropy (CE) Loss in existing frameworks . The formulation of CE Loss can be written as
where and are the predicted logits and probabilities of the classifier, respectively. And is the one-hot ground truth label. Given a training sample of class , the gradients on and are given by
It shows that samples of class punish the classifier of class w.r.t. . In the case that the instance number of class is enormously greater than that of class , the classifier of class will receive penalties in most samples and attains few positive signals during training. Thus the predicted probabilities of class will be heavily suppressed, which results in a low classification accuracy of tail classes, as shown in Figure 1.
2 Seesaw Loss
To alleviate the above mentioned problem, one feasible solution is to decrease the gradients of negative samples in Eq. 3 imposed by head classes on a tail class. Therefore, we propose Seesaw Loss as
Then the gradient on of negative class in Eqn 3 becomes
Here works as a tunable balancing factor between different classes. By a careful design of , Seesaw loss adjusts the punishments on class from positive samples of class . Seesaw loss determines by a mitigation factor and a compensation factor, as
The mitigation factor decreases the penalty on tail class according to a ratio of instance numbers between tail class and head class . The compensation factor increases the penalty on class whenever an instance of class is misclassified to class .
Mitigation Factor. Seesaw Loss accumulates instance number for each category at each iteration in the whole training process. As shown in Fig. 2, given an instance with positive label , for another category , the mitigation factor adjusts the penalty for negative label w.r.t. the ratio
When category is more frequent than category , Seesaw Loss will reduce the penalty on category , which is imposed by samples of category , by a factor of . Otherwise, Seesaw Loss will keep the penalty on negative classes to reduce misclassification. The exponent is a hyper-parameter that adapts the magnitude of mitigation.
Note that Seesaw Loss accumulates the instance numbers during training, rather than get the statistics from the whole dataset ahead of time. This strategy brings two benefits. First, it can be applied when the distribution of the whole training set is unavailable, e.g., training examples are obtained from a stream. Second, the training samples of each category can be affected by the adopted data sampler , and the online accumulation is robust to sampling methods. During training, the mitigation factor is uniformly initialized and smoothly updated to approximate the real data distribution.
Compensation Factor. The mitigation factor effectively balances the gradients of head and tail classes. Nevertheless, it may cause more false positives for tail classes due to less penalty. Moreover, the false positives cannot be eliminated by simply adjusting in , since it is applied to the whole category. We propose a compensation factor that focuses on misclassified samples instead of adjusting the whole category. As shown in Fig. 2, this factor compensates the diminished gradient when there is misclassification, i.e., the predicted probability of negative label is greater than . The compensation factor is calculated as
For a training sample with positive label , if the predicted probability of any negative class is greater than class , i.e., , the compensation factor increases the punishment on class by a factor of , where is a hyper-parameter to control the scale. Otherwise, and only the mitigation factor is applied.
Normalized Linear Activation. The classifier in an object detector usually predicts classification logits as on the dataset with balanced category distribution , where and are the weights and bias of the linear layer and is the input features. On long-tailed datasets, previous works find that the weight norm of is highly related to the number of training instances in the corresponding category . The more training samples of category there are, the larger will be. This phenomenon is also observed in the feature norm . Therefore, we adopt a normalized linear activation (which are related) to re-balance the scale of and as , where , , and is a temperature factor. The normalized linear activation normalizes the weights and features by their norm to reduce their scale variance for different categories. Thus, it effectively balances the distribution of predicted probabilities of different categories and improves the performance on a long-tailed dataset.
3 Model Design for Instance Segmentation
Objectness Branch. In contrast to image classification, the classifier in an object detector has two functionalities. It first determines if a bounding box is a foreground object then distinguishes which category the foreground instance belongs to. Previous practices usually regard the background as an auxiliary category in the classifier. Given a dataset with categories, the classifier in most detectors predicts logits of classes. Although widely adopted, this design brings difficulty when adopting Seesaw Loss to balance long-tailed distribution. In general, most object candidates in a detector are backgrounds. Thus, all foreground categories are much rarer categories compared to the background category. Consequently, Seesaw Loss will significantly reduce punishments on all foreground categories. As a result, the classifier tends to misclassify more backgrounds as foregrounds and harms the performance.
To tackle this problem, we decouple the two functionalities of the classifier in an object detecter. Specifically, apart from the classifier with classes, we adopt an extra objectness branch to distinguish the foregrounds and backgrounds. The objectness branch adopts the normalized linear activation to predict logits of two classes, i.e., foreground and background, and is trained by cross-entropy loss. During inference, both the classification logit of category and logit of objectness are activated with a softmax function. The final detection probability for a bounding box of category is .
Normalized Mask Predication. Inspired by normalized linear activation, we further present a normalized mask prediction to alleviate the biased training process in mask head. In Mask R-CNN , a 1x1 convolution layer is applied in the end of the mask head, and the predicted logits are activated by a sigmoid function. We normalize the weights of the 1x1 convolution layer and the input features with normalization. Note that the spatial size of is , we denote the feature at as . The formula of normalized mask prediction is , where , and is a temperature factor.
Experiments
Datasets. We perform experiments on the challenging LVIS v1 dataset . LVIS is a large vocabulary instance segmentation dataset containing 1203 categories with high-quality instance mask annotations. LVIS v1 provides a train split with 100k images, a val split with 19.8k images and a test-dev split with 19.8k images. According to the numbers of images that each category appears in the train split, the categories are divided into three groups: rare (1-10 images), common (11-100 images) and frequent (100 images).
Evaluation metrics. The results of instance segmentation are evaluated with of mask prediction, which is averaged at different IoU thresholds (from 0.5 to 0.95) across categories. The AP for rare, common and frequent categories are denoted as , and . The AP for detection boxes is denoted as .
Implementation Details. We implement our method with mmdetection and train Mask R-CNN , Cascade Mask R-CNN using the 2x training schedule . The model is trained with batch size of 16 for 24 epochs. The learning rate is 0.02, and it will decrease by 0.1 after 16 and 22 epochs, respectively. ResNet-50 with FPN backbone is adopted if not further specified. Following the practice in mmdetection , we adopt multi-scale with horizontally flip augmentation during training. Specifically, we randomly resize the shorter edge of the image within 640, 672, 704, 736, 768, 800 pixels and keep the longer edge smaller than 1333 pixels without changing the aspect ratio. In inference, we adopt single-scale testing with image size of pixels and score thresholds of without bells and whistles.
Apart from the standard random sampler that samples images in train split randomly, the repeat factor sampler (RFS) is also evaluated in experiments. RFS oversamples categories that appear in less than 0.1% of the total images and is effective to improve the overall AP. The ablation study is conducted with RFS if not further specified. We adopt Seesaw Loss in the box classification branch of Mask R-CNN with hyper-parameter , , and . In Cascade Mask R-CNN , Seesaw Loss is adopted in box classification branches of all three stages with the same hyper-parameters as that in Mask R-CNN. We further evaluate the proposed Normalized Mask Prediction and integrate it into the mask head of Mask R-CNN and all the mask heads in Cascade Mask R-CNN. For simplicity, Normalized Mask Prediction adopts the same temperature, i.e. . We use the train split for training and report the performance on val split for ablation study. The performance of our method is also reported on test-dev split.
2 Benchmark Results
To show the effectiveness of Seesaw Loss, we perform extensive experiments with different data samplers and instance segmentation frameworks. We adopt Mask R-CNN with ResNet-101 backbone with FPN and train the models with the random sampler or the repeat factor sampler (RFS) by 2x schedule.
As shown in Table 1, Seesaw Loss significantly outperforms Cross-Entropy (CE) loss by 6.0% AP with random sampler and 2.1% AP on the stronger baseline with RFS. The improvements on , , and with both samplers reveals the effectiveness of Seesaw Loss on categories with different frequency. We further integrate the proposed Normalized Mask Prediction (Norm Mask) into Mask R-CNN with Seesaw Loss. Without extra cost, the overall AP is improved from 26.6% to 27.1 % and 27.6% to 28.1% with random sampler and RFS, respectively.
Apart from the CE loss baseline, we further compare Seesaw Loss with recent designs for long-tailed instance segmentation, i.e., Equalization Loss (EQL) and Balanced Group Softmax (BAGS) , in Table 1. Seesaw Loss outperforms EQL by 3.9% AP and 1.4% AP, and outperforms BAGS by 1.0% AP and 1.8% AP with random sampler and RFS, respectively. Seesaw Loss also achieves higher , and than the two methods consistently. Notably, EQL and BAGS achieve lower than the CE baseline while Seesaw Loss does not. This phenomenon indicates that these two methods improve the performance of rare and common categories while sacrificing frequent categories.
We further compare Seesaw Loss with previous methods with both random sampler and RFS on Cascade Mask R-CNN . It’s a representative framework of cascade methods that outperforms Mask R-CNN . As shown in Table 1, Seesaw Loss performs much superior to previous works on Cascade Mask R-CNN. Specifically, Seesaw Loss improves the baseline by 6.4% AP and 2.3% AP with random sampler and RFS, respectively. With Normalized Mask Prediction, Cascade Mask-RCNN with Seesaw Loss finally achieves 29.6% AP and 30.1% AP with the two samplers, respectively. Moreover, Seesaw Loss is also evaluated on test-dev split and consistently obtains significant gains over the CE baseline.
3 Ablation study
We conduct a comprehensive ablation study to verify the effectiveness of each design choice in the proposed method.
Components in Seesaw Loss. There are three components in Seesaw Loss: mitigation factor, compensation factor, and normalized linear activation. We evaluated each component on Mask R-CNN with RFS (Table 2). The mitigation factor that mitigates the overwhelming punishments on rare classes leads to a significant improvement from 23.7% AP to 25.1% AP. Notably, it improves the of rare classes by 2.8% AP. The compensation factor increases the punishments of a class when it observes false positives on that class to reduce misclassification. It improves the baseline by 0.4% AP. The combination of the mitigation and the compensation factors achieves 25.7% AP, outperforming the performance of mitigation factor by 0.6% AP. It reveals the effectiveness of instance-wise compensation to avoid misclassification. The normalized linear activation is another important component in Seesaw Loss, which reduces the scale invariance of weights and features across different categories. It improves the baseline performance from 23.7% to 24.7% AP. Seesaw Loss combining all these three components achieves 26.4% AP.
Normalized linear activation. We empirically find that normalized linear activation (NLA) helps to improve the performance of both CE Loss and Seesaw Loss. Therefore, we further integrate NLA with equalization loss (EQL) and balanced group softmax (BAGS) for fair comparisons. Results in Table 3 show that NLA improves the performance of EQL and BAGS by 0.3% and 0.8% AP, respectively. It is noteworthy that Seesaw Loss outperforms EQL and BAGS no matter whether NLA is adopted.
Cumulative Sample Numbers. Different from previous works that rely on the pre-computed frequency distribution of categories in the dataset, Seesaw Loss accumulates the sample numbers of each category during training. We compare different approaches to obtain the sample numbers of categories for Seesaw Loss (Table 4). Directly using the statistics of train split decreases the performance of Seesaw Loss by 0.3% AP. The reason lies in that the data sampler, e.g., repeat factor sampler, changes the frequency distribution of categories during training. We also explore loading the pre-recorded distribution of training samples from a model trained with Seesaw Loss. It achieves a similar performance with accumulating the training samples online (26.3% AP \vs26.4% AP). These results verify the effectiveness and simplicity of online accumulating.
Hyper-parameters. We study the hyper-parameters, i.e., , , , adopted in different components of Seesaw Loss. The normalized linear activation is not applied when studying the mitigation and compensation factors (25.7% AP with this setting). In Table 5, we explore in of mitigation factor. controls the magnitude to mitigate the punishments on rare classes. A higher value of will reduce punishments more, as well as increase the risk of inducing false positives of tail classes. Therefore, it is critical to find a suitable . Results show that achieves the best performance. In Table 6, we explore in of the compensation factor. controls the magnitude to compensate the reduced punishments on tail classes when false positives are observed. We study the effectiveness of different and find achieves the best performance. Notably, is robust across different values as the best value is only 0.3% AP better than the worst value. In Table 7, we study the temperature in normalized linear activation (NLA). determines the variance of the classifier’s predicted logit . If is too small, the variance of is insufficient to distinguish positive and negative samples. However, if is too big, the target of balancing the variance in weights and features between different categories will be sacrificed. We choose in the NLA as it achieves the best performance.
Objectness Branch. In a common practice of object detection, the classifier predicts scores for foregrounds categories and one background category. Due to the extremely imbalanced distribution between foregrounds and backgrounds, Seesaw Loss will tend to misclassify more backgrounds as foregrounds with this design. Thus, we adopt an extra objectness branch as described in Sec. 3.3. The results in Table 8 shows that the objectness branch does not improve the Coss-Entropy loss baseline but is critical to Seesaw Loss. The objectness branch helps to avoid reducing the backgrounds’ punishments on foreground categories. As a result, the objectness branch brings gains on Seesaw Loss across categories with different frequency, and improves the overall AP from 25.3% to 26.4%.
Training Pipeline. Apart from the end-to-end training pipeline, we further explore the popular decoupling training pipeline on Mask R-CNN. Specifically, we pre-train the Mask R-CNN with Cross-Entropy loss using either random sampler or repeat factor sampler for 2x schedule. Then we finetune the final fully-connected layer of the classifier with all other components fixed. The 1x schedule and repeat factor sampler is adopted during finetuning. As shown in Table 9, Seesaw Loss with the pre-trained model on repeat factor sampler (P-RFS) achieves 25.8% AP, outperforming other methods with decoupling training pipeline. Notably, Seesaw Loss performs better with the end-to-end training pipeline than with the decoupling training pipeline. It indicates Seesaw Loss provides a simpler and more effective solution to long-tailed instance segmentation without relying on complex training pipelines.
4 Long-Tailed Image Classification
To show the versatility of Seesaw Loss, we apply it for long-tailed image classification task on ImageNet-LT dataset. ImageNet-LT is generated from the ImageNet-2012 dataset with long-tailed distributed categories in training set. There are 115.8k images of 1000 categories with a maximum number of 1280 images and a minimum number of 5 images. The performance is evaluated with top-1 accuracy on all categories and the accuracies for Many Shot ( 100 images), Medium Shot (20100 images) and Few Shot ( 20 images) categories are also reported.
We adopt two training pipelines: end-to-end training and decoupling training. We use ResNeXt-50 backbone and SGD optimizer with momentum of 0.9, initial learning rate of 0.2, batch size of 512, and cosine learning rate following . For the end-to-end training pipeline, the model is trained for 90 epochs. For decouple training pipeline , we load the pre-trained ResNeXt-50 with Cross-Entropy Loss (CE), and finetune the classifier with class-balanced sampler while fixing all other layers for 10 epochs. Seesaw Loss in image classification mostly follows the hyper-parameters on the instance segmentation task except for in the compensation factor. We adopt for ImageNet-LT dataset. The study of for ImageNet-LT dataset is shown in Table 11.
We report the performance of Seesaw Loss in Table 10. Seesaw Loss improves top-1 accuracies of CE from 44.4% to 49.7% and 50.4% with the decoupling training and the end-to-end training pipeline, respectively. Similar to our observations on the instance segmentation task, Seesaw Loss performs better with the end-to-end training pipeline on image classification. The performance achieved by Seesaw Loss with the end-to-end pipeline is competitive among previous methods on ImageNet-LT .
Conclusion
In this paper, we propose Seesaw Loss for long-tailed instance segmentation. Seesaw Loss dynamically re-balances gradients of positive and negative samples for each category with two complementary factors. The mitigation factor reduces punishments to tail categories w.r.t. the ratio of cumulative training instances between categories. Meanwhile, the compensation factor increases the penalty of misclassified instances to avoid false positives. Experimental results demonstrate that Seesaw Loss provides a simpler and more effective solution to long-tailed instance segmentation without relying on complex training pipelines.
Acknowledgements. This research was conducted in collaboration with SenseTime. This work is supported by GRF 14203518, ITS/431/18FX, CUHK Agreement TS1712093, NTU NAP, A*STAR through the Industry Alignment Fund - Industry Collaboration Projects Grant, Shanghai Committee of Science and Technology, China (Grant No. 20DZ1100800).
In this work, we propose Seesaw Loss to dynamically re-balance gradients of positive and negative samples for each category. Specifically, Seesaw Loss mitigates the overwhelming gradients of negative samples imposed by a head class on a tail class via decreasing the value of in the following formula,
To further analyze the effects of adjusting the value of , we calculate the partial derivative of Eqn 9 with respect to as
The value of the partial derivative in Eqn 10 is always positive. This indicates that the gradients of negative samples imposed by class on class will be reduced as the value of decreases.
Appendix B How Seesaw Loss works
Via re-balancing gradients of positive and negative samples, Mask R-CNN w/ Seesaw Loss significantly outperforms Mask R-CNN w/ Cross-Entropy Loss on LVIS dataset. Here, we conduct a quantitative analysis of the effectiveness of Seesaw Loss on re-balancing the gradients of positive and negative samples for each category. Specifically, we adopt Mask R-CNN with ResNet-101 backbone and FPN as instance segmentation framework. The Cross-Entropy Loss and Seesaw Loss are integrated into the framework and trained with random sampler by 2x schedule. We accumulate the gradients of positive and negative samples on predicted logit of each category during the whole training procedure.
Figure 3 shows the distribution of the ratio of cumulative gradients between positive and negative samples for each category in Mask R-CNN with Cross-Entropy Loss and Seesaw Loss, respectively. With Cross-Entropy Loss, tail classes obtain heavily imbalanced gradients of positive and negative samples during training. The overwhelming gradients of negative samples lead to a biased learning process for the classifier, which results in the low classification accuracy on tail classes. On the contrary, Seesaw Loss effectively re-balances the gradients of positive and negative samples across different categories. Consequently, Mask R-CNN with Seesaw Loss achieves significant improvements on instance segmentation performance as shown in Figure 1 and Table 1 in the main text.
Appendix C Per-category Performance Comparison
In addition to the performance reported in Table 1 of the main text, we further show the per-category performance (AP) to verify the superiority of Seesaw Loss compared to other loss functions. As shown in Figure 4, compared to other loss functions (i.e., Cross-Entropy Loss, Equalization Loss , and Balanced Group Softmax ), Seesaw Loss consistently achieves strong performance across categories with different frequency on different frameworks (i.e. Mask R-CNN , Cascade Mask R-CNN ) and samplers (i.e., random sampler, repeat factor sampler ).
Appendix D LVIS Challenge 2020
Here we present the approach used in the entry of team MMDet in the LVIS Challenge 2020. In our entry, we adopt Seesaw Loss for long-tailed instance segmentation as described in the main text. Seesaw Loss improves the strong baseline by 6.9% AP on LVIS v1 val split. Furthermore, we propose HTC-Lite, a light-weight version of Hybrid Task Cascade (HTC) which replaces the semantic segmentation branch with a global context encoder. With a single model and without using external data and annotations except for standard ImageNet-1k classification dataset for backbone pre-training, our entry achieves 38.92% AP on the test-dev split of the LVIS v1 benchmark.
We propose HTC-Lite, a light-weight version of Hybrid Task Cascade (HTC) , to accelerate the training and inference speed while maintaining good performance. As shown in Figure 5, the modifications are in two folds: replacing the semantic segmentation branch with a global context encoding branch and reducing mask heads.
Context Encoding Branch. Since semantic segmentation annotations are not available for LVIS dataset, we replace the semantic segmentation branch with a global context encoder which works as a multi-label classification branch trained by a binary cross-entropy loss. The context encoder applies convolution layers and a global average pooling on the input feature map to obtain a feature vector. And an auxiliary fully connected (fc) layer is applied on the feature vector to predict the categories existing in the current image. By this approach, this feature vector encodes the global context information of the image. Then it is added to the RoI features used by box heads and mask heads to enrich their semantic information.
Reduced Mask Heads. To further reduce the cost of instance segmentation, HTC-Lite only keeps the mask head in the last stage, which also spares the original interleaved information passing.
In Table 13, we compare the performance and inference speed on LVIS v1 dataset of HTC-Lite with two mainstream cascading instance segmentation frameworks, i.e., Cascade Mask R-CNN and HTC. The ResNet-50 with FPN backbone, repeat factor sampler and 1x training schedule are adopted in these methods. The semantic segmentation branch in HTC is removed since semantic segmentation annotations are not available on LVIS v1 dataset. We evaluate the inference speed for each framework with a single Tesla V100 GPU. The experimental results show that HTC-Lite is not only much more efficient than its counterparts but also outperforms them.
D.2 Step by Step Results
We report the step-by-step results of our entry in LVIS Challenge 2020 as shown in Table 12.
Baseline. The baseline model is Mask R-CNN using ResNet-50-FPN , trained with multi-scale training and random data sampler by 2x schedule .
SyncBN. We use SyncBN in the backbone and heads.
CARAFE Upsample. CARAFE is used for upsampling in the mask head.
HTC-Lite. We use HTC-Lite as described in Appendix D.1.
TSD. TSD is used to replace the box heads in all three stages in HTC-Lite.
Mask Scoring. We further use the mask IoU head to improve mask results.
Training Time Augmentation. We train the model with stronger augmentations with 45 epochs. The learning rate is decreased by 0.1 at 30 and 40 epochs. We randomly resize the image with its longer edge in a range of 768 to 1792 pixels. And then, we randomly crop the image to the size of after adopting instaboost augmentation .
Stronger Neck. We replace the neck architecture with an enhanced version of Feature Pyramid Grids (FPG) . The enhanced FPG uses deformable convolution v2 (DCNv2) after feature upsampling, and a downsampler version of CARAFE for feature downsampling.
Stronger Backbone. We use ResNeSt-200 with DCNv2 as the backbone.
Seesaw Loss. We apply the proposed Seesaw Loss to classification branches of the TSD box head, in all cascading stages. Furthermore, we remove the original progressive constraint (PC) loss on classification branches in TSD.
Dual Head Classification. Inspired by , we adopt a dual-head classification policy to further boost the performance. Specifically, after obtaining the model with Seesaw Loss trained by a random sampler, we freeze all components in the original model. Then we finetune a new classification branch for each cascading stage on the fixed model using repeat factor sampler by 1x schedule. During inference, the classification scores of original classification branches and the scores of new classification branches are averaged to get the final scores.
Test Time Augmentation. We adopt multi-scale testing with horizontal flipping. Specifically, image scales are 1200, 1400, 1600, 1800, and 2000 pixels.
Final Performance on Test-dev. After adding the abovementioned components step by step, we finally achieve 38.8% AP on the val split and 38.92% AP on the test-dev split.