Seesaw Loss for Long-Tailed Instance Segmentation

Jiaqi Wang, Wenwei Zhang, Yuhang Zang, Yuhang Cao, Jiangmiao Pang, Tao Gong, Kai Chen, Ziwei Liu, Chen Change Loy, Dahua Lin

Introduction

Deep learning-based object detection and instance segmentation approaches have achieved immense success on datasets with relatively balanced category distribution, e.g., COCO dataset . However, the distribution of categories in the real world is long-tailed . There are a few head classes containing abundant instances, while most other classes comprise relatively few instances.

On long-tailed datasets, existing instance segmentation frameworks fail to perform as accurately as on the datasets with balanced category distribution, exhibiting unsatisfactory performance on tail classes. Figure 1 shows the classification accuracy and instance segmentation performance of Mask R-CNN on LVIS dataset. The classifier in Mask R-CNN trained by Cross-Entropy Loss tends to misclassify tail categories as backgrounds or other confusing head classes, which leads to extremely low accuracy on tail classes.

The primary reason for this undesired phenomenon is that the instances from head classes are predominant in a long-tailed dataset. These instances contribute an overwhelmingly large quantity of negative samples for tail classes. Thus, the gradients of positive and negative samples on a tail class are heavily imbalanced, leading to a biased learning process for the classifier. One can imagine that gradients of positive and negative samples resemble two objects positioned on each end of a seesaw (see Fig. 1). To balance them, a viable solution is to shorten the arm of the heavier end in the seesaw, which is equivalent to scaling down the overwhelming gradients of negative samples on the tail class by a factor. Nevertheless, blindly reducing the gradients of negative samples increases the risk of inducing false positives of tail classes, since samples of other classes are less punished when they are misclassified as tail classes. Thus, a specialized mechanism is needed to compensate for the excessively reduced penalties on tail classes.

In this work, we propose Seesaw Loss that dynamically re-balances positive and negative gradients for each category with two complementary factors, i.e., mitigation factor and compensation factor. According to the ratio between categories’ cumulative sample numbers during training, the mitigation factor reduces the penalty to relatively rare classes. When a false positive sample of one category is observed, the compensation factor will increase the penalty to that category. The synergy of the two above factors enables Seesaw Loss to mitigate the overwhelming punishments to tail classes as well as compensate for the risk of misclassification caused by diminished penalties.

Seesaw Loss has three appealing properties. 1) Seesaw Loss is dynamic. It explores the ratios of cumulative training sample numbers between different categories and instance-wise misclassification during training. This differs significantly to previous solutions that rely either on static group split or loss reweighting with constant values . 2) Seesaw Loss is self-calibrated. The mitigation and the compensation factor synergize to relieve the overwhelming punishments on tail classes as well as avoid increasing false positives of tail categories. On the contrary, previous methods blindly reduce punishments on tail classes or decrease the loss weights of head categories . 3) Seesaw Loss is distribution-agnostic. It does not rely on pre-computed datasets’ distribution , and it can operate well with any data sampler . By accumulating the number of samples in each class, Seesaw Loss gradually approximates the real data distribution during training to achieve more accurate balancing.

Through extensive experiments, we show consistent improvements of Seesaw Loss in different instance segmentation frameworks and data samplers. On the challenging LVIS dataset, Seesaw Loss achieves significant improvements of 6.0% AP and 2.1% AP upon Mask R-CNN with random sampler and repeat factor sampler , respectively. Even if switching to the stronger Cascade Mask R-CNN , we still observe an impressive improvement of 6.4% AP and 2.3% AP with random sampler and repeat factor sampler. To show the versatility of Seesaw Loss, we integrate it into the long-tailed image classification task. Seesaw Loss significantly improves the classification accuracy by 6% on ImageNet-LT dataset. Besides, we also explore the necessity of the decoupling training pipeline in Seesaw Loss. Experimental results demonstrate that Seesaw Loss provides a simpler and more effective solution to long-tailed instance segmentation without relying on complex training pipelines.

Related Work

Object Detection. Recent years have witnessed a remarkable improvement in object detection . A leading paradigm in this area is the two-stage pipeline , where the first stage generates a set of region proposals, and then the second stage classifies and refines the proposals. Unlike the two-stage approaches, the single-stage pipeline directly predicts bounding boxes. Classical single-stage approaches require densely populated anchors as a prior, while anchor-free methods manage to achieve similar or better performance without such prior. There are also attempts to apply cascade architecture to refine the bounding boxes’ predictions progressively.

Instance Segmentation. Instance segmentation is becoming popular in tandem with a surge in the interest in object detection. Early methods perform segmentation before object recognition . Via adding a mask prediction branch in the Faster R-CNN architecture, Mask R-CNN bridges the gap between object detection and instance segmentation. The idea is also adopted by in their cascading frameworks. More recent works introduce an even shorter pipeline by skipping the detection process and directly predicting mask for each instance. Seesaw Loss can easily cooperates with object detection and instance segmentation frameworks for the long-tailed datasets.

Long-Tailed Recognition. Long-tailed recognition tasks receive growing attention recently as the problems are closer to real-world applications. One representative solution to the problem is loss re-weighting . Loss re-weighting methods adopt different re-weighting strategies to adjust the loss of different classes based on each class’s statistics . Other common approches re-balance the distribution of the instance numbers in each class, e.g., repeat factor sampling and class-balanced sampling , both are based on the sample numbers of classes. Different sampling strategies can be adopted at different training stages to formulate a multi-stage training procedure . A recent work proposes a decoupling training pipeline. which first trains a good representation network with natural sampling and then finetunes the classifier with class-balanced sampling. There are also attempts to modify the classifier to improve the performance on tail classes, e.g., using different classifiers for different groups of classes , or use two classifiers trained with different data samplers .

Methodology

The classifier trained by the widely applied Cross-Entropy (CE) Loss (Sec. 3.1) is highly biased on long-tailed datasets, resulting in much lower accuracy of tail classes than head classes. The major reason is that gradients brought by positive samples are overwhelmed by gradients from negative samples on tail classes. Therefore, we propose Seesaw Loss to mitigate the overwhelming gradients of negative samples on tail classes as well as compensate the gradients of misclassified samples to avoid false positives (Sec. 3.2). We also explore some practical component designs to adopt Seesaw Loss in instance segmentation (Sec. 3.3).

We first revisit the most widely adopted Cross-Entropy (CE) Loss in existing frameworks . The formulation of CE Loss can be written as

where z=[z1,z2,…,zC]\mathbf{z}=[z_{1},z_{2},\dots,z_{C}] and σ=[σ1,σ2,…,σC]\boldsymbol{\sigma}=[\sigma_{1},\sigma_{2},\dots,\sigma_{C}] are the predicted logits and probabilities of the classifier, respectively. And yi∈{0,1},1≤i≤Cy_{i}\in\{0,1\},1\leq i\leq C is the one-hot ground truth label. Given a training sample of class ii, the gradients on ziz_{i} and zjz_{j} are given by

It shows that samples of class ii punish the classifier of class jj w.r.t. σj\sigma_{j}. In the case that the instance number of class ii is enormously greater than that of class jj, the classifier of class jj will receive penalties in most samples and attains few positive signals during training. Thus the predicted probabilities of class jj will be heavily suppressed, which results in a low classification accuracy of tail classes, as shown in Figure 1.

2 Seesaw Loss

To alleviate the above mentioned problem, one feasible solution is to decrease the gradients of negative samples in Eq. 3 imposed by head classes on a tail class. Therefore, we propose Seesaw Loss as

Then the gradient on zjz_{j} of negative class jj in Eqn 3 becomes

Here Sij\mathcal{S}_{ij} works as a tunable balancing factor between different classes. By a careful design of Sij\mathcal{S}_{ij}, Seesaw loss adjusts the punishments on class jj from positive samples of class ii. Seesaw loss determines Sij\mathcal{S}_{ij} by a mitigation factor and a compensation factor, as

The mitigation factor Mij\mathcal{M}_{ij} decreases the penalty on tail class jj according to a ratio of instance numbers between tail class jj and head class ii. The compensation factor Cij\mathcal{C}_{ij} increases the penalty on class jj whenever an instance of class ii is misclassified to class jj.

Mitigation Factor. Seesaw Loss accumulates instance number NiN_{i} for each category ii at each iteration in the whole training process. As shown in Fig. 2, given an instance with positive label ii, for another category jj, the mitigation factor adjusts the penalty for negative label jj w.r.t. the ratio NjNi\frac{N_{j}}{N_{i}}

When category ii is more frequent than category jj, Seesaw Loss will reduce the penalty on category jj, which is imposed by samples of category ii, by a factor of (NjNi)p\left(\frac{N_{j}}{N_{i}}\right)^{p}. Otherwise, Seesaw Loss will keep the penalty on negative classes to reduce misclassification. The exponent pp is a hyper-parameter that adapts the magnitude of mitigation.

Note that Seesaw Loss accumulates the instance numbers during training, rather than get the statistics from the whole dataset ahead of time. This strategy brings two benefits. First, it can be applied when the distribution of the whole training set is unavailable, e.g., training examples are obtained from a stream. Second, the training samples of each category can be affected by the adopted data sampler , and the online accumulation is robust to sampling methods. During training, the mitigation factor is uniformly initialized and smoothly updated to approximate the real data distribution.

Compensation Factor. The mitigation factor effectively balances the gradients of head and tail classes. Nevertheless, it may cause more false positives for tail classes due to less penalty. Moreover, the false positives cannot be eliminated by simply adjusting pp in Mij\mathcal{M}_{ij}, since it is applied to the whole category. We propose a compensation factor that focuses on misclassified samples instead of adjusting the whole category. As shown in Fig. 2, this factor compensates the diminished gradient when there is misclassification, i.e., the predicted probability σj\sigma_{j} of negative label jj is greater than σi\sigma_{i}. The compensation factor Cij\mathcal{C}_{ij} is calculated as

For a training sample with positive label ii, if the predicted probability of any negative class jj is greater than class ii, i.e., σj>σi\sigma_{j}>\sigma_{i}, the compensation factor increases the punishment on class jj by a factor of (σjσi)q\left(\frac{\sigma_{j}}{\sigma_{i}}\right)^{q}, where qq is a hyper-parameter to control the scale. Otherwise, Cij=1\mathcal{C}_{ij}=1 and only the mitigation factor Mij\mathcal{M}_{ij} is applied.

Normalized Linear Activation. The classifier in an object detector usually predicts classification logits as z=WTx+bz=\mathcal{W}^{T}x+b on the dataset with balanced category distribution , where W\mathcal{W} and bb are the weights and bias of the linear layer and xx is the input features. On long-tailed datasets, previous works find that the weight norm of Wi\mathcal{W}_{i} is highly related to the number of training instances in the corresponding category ii. The more training samples of category ii there are, the larger ∥Wi∥\left\|\mathcal{W}_{i}\right\| will be. This phenomenon is also observed in the feature norm ∥x∥\left\|x\right\|. Therefore, we adopt a normalized linear activation (which are related) to re-balance the scale of ∥Wi∥\left\|\mathcal{W}_{i}\right\| and ∥x∥\left\|x\right\| as z=τW~Tx~+bz=\tau{\widetilde{\mathcal{W}}}^{T}\widetilde{x}+b, where W~i=Wi∥Wi∥2,i∈C\widetilde{\mathcal{W}}_{i}=\frac{\mathcal{W}_{i}}{\left\|\mathcal{W}_{i}\right\|_{2}},i\in C, x~=x∥x∥2\widetilde{x}=\frac{x}{\left\|x\right\|_{2}}, and τ\tau is a temperature factor. The normalized linear activation normalizes the weights W\mathcal{W} and features xx by their L2L_{2} norm to reduce their scale variance for different categories. Thus, it effectively balances the distribution of predicted probabilities of different categories and improves the performance on a long-tailed dataset.

3 Model Design for Instance Segmentation

Objectness Branch. In contrast to image classification, the classifier in an object detector has two functionalities. It first determines if a bounding box is a foreground object then distinguishes which category the foreground instance belongs to. Previous practices usually regard the background as an auxiliary category in the classifier. Given a dataset with CC categories, the classifier in most detectors predicts logits of C+1C+1 classes. Although widely adopted, this design brings difficulty when adopting Seesaw Loss to balance long-tailed distribution. In general, most object candidates in a detector are backgrounds. Thus, all foreground categories are much rarer categories compared to the background category. Consequently, Seesaw Loss will significantly reduce punishments on all foreground categories. As a result, the classifier tends to misclassify more backgrounds as foregrounds and harms the performance.

To tackle this problem, we decouple the two functionalities of the classifier in an object detecter. Specifically, apart from the classifier with CC classes, we adopt an extra objectness branch to distinguish the foregrounds and backgrounds. The objectness branch adopts the normalized linear activation to predict logits of two classes, i.e., foreground and background, and is trained by cross-entropy loss. During inference, both the classification logit ziclassz^{class}_{i} of category i∈Ci\in C and logit of objectness zobjz^{obj} are activated with a softmax function. The final detection probability σidet\sigma^{det}_{i} for a bounding box of category ii is σidet=σiclass⋅σobj\sigma^{det}_{i}=\sigma^{class}_{i}\cdot\sigma^{obj}.

Normalized Mask Predication. Inspired by normalized linear activation, we further present a normalized mask prediction to alleviate the biased training process in mask head. In Mask R-CNN , a 1x1 convolution layer is applied in the end of the mask head, and the predicted logits are activated by a sigmoid function. We normalize the weights W\mathcal{W} of the 1x1 convolution layer and the input features X\mathcal{X} with L2L2 normalization. Note that the spatial size of X\mathcal{X} is H×WH\times W, we denote the feature at (y,x)(y,x) as Xy,x\mathcal{X}_{y,x}. The formula of normalized mask prediction is z=τW~∗X~+bz=\tau\widetilde{\mathcal{W}}\ast\widetilde{\mathcal{X}}+b, where W~i=Wi∥Wi∥2,i∈C\widetilde{\mathcal{W}}_{i}=\frac{\mathcal{W}_{i}}{\left\|\mathcal{W}_{i}\right\|_{2}},i\in C, X~y,x=Xy,x∥Xy,x∥2,y∈H,x∈W\widetilde{\mathcal{X}}_{y,x}=\frac{\mathcal{X}_{y,x}}{\left\|\mathcal{X}_{y,x}\right\|_{2}},y\in H,x\in W and τ\tau is a temperature factor.

Experiments

Datasets. We perform experiments on the challenging LVIS v1 dataset . LVIS is a large vocabulary instance segmentation dataset containing 1203 categories with high-quality instance mask annotations. LVIS v1 provides a train split with 100k images, a val split with 19.8k images and a test-dev split with 19.8k images. According to the numbers of images that each category appears in the train split, the categories are divided into three groups: rare (1-10 images), common (11-100 images) and frequent (>>100 images).

Evaluation metrics. The results of instance segmentation are evaluated with APAP of mask prediction, which is averaged at different IoU thresholds (from 0.5 to 0.95) across categories. The AP for rare, common and frequent categories are denoted as APr\text{AP}_{r}, APc\text{AP}_{c} and APf\text{AP}_{f}. The AP for detection boxes is denoted as APbox\text{AP}^{box}.

Implementation Details. We implement our method with mmdetection and train Mask R-CNN , Cascade Mask R-CNN using the 2x training schedule . The model is trained with batch size of 16 for 24 epochs. The learning rate is 0.02, and it will decrease by 0.1 after 16 and 22 epochs, respectively. ResNet-50 with FPN backbone is adopted if not further specified. Following the practice in mmdetection , we adopt multi-scale with horizontally flip augmentation during training. Specifically, we randomly resize the shorter edge of the image within {\{640, 672, 704, 736, 768, 800}\} pixels and keep the longer edge smaller than 1333 pixels without changing the aspect ratio. In inference, we adopt single-scale testing with image size of 1333×8001333\times 800 pixels and score thresholds of 10−310^{-3} without bells and whistles.

Apart from the standard random sampler that samples images in train split randomly, the repeat factor sampler (RFS) is also evaluated in experiments. RFS oversamples categories that appear in less than 0.1% of the total images and is effective to improve the overall AP. The ablation study is conducted with RFS if not further specified. We adopt Seesaw Loss in the box classification branch of Mask R-CNN with hyper-parameter p=0.8p=0.8, q=2q=2, and τ=20\tau=20. In Cascade Mask R-CNN , Seesaw Loss is adopted in box classification branches of all three stages with the same hyper-parameters as that in Mask R-CNN. We further evaluate the proposed Normalized Mask Prediction and integrate it into the mask head of Mask R-CNN and all the mask heads in Cascade Mask R-CNN. For simplicity, Normalized Mask Prediction adopts the same temperature, i.e. τ=20\tau=20. We use the train split for training and report the performance on val split for ablation study. The performance of our method is also reported on test-dev split.

2 Benchmark Results

To show the effectiveness of Seesaw Loss, we perform extensive experiments with different data samplers and instance segmentation frameworks. We adopt Mask R-CNN with ResNet-101 backbone with FPN and train the models with the random sampler or the repeat factor sampler (RFS) by 2x schedule.

As shown in Table 1, Seesaw Loss significantly outperforms Cross-Entropy (CE) loss by 6.0% AP with random sampler and 2.1% AP on the stronger baseline with RFS. The improvements on APr\text{AP}_{r}, APc\text{AP}_{c}, and APf\text{AP}_{f} with both samplers reveals the effectiveness of Seesaw Loss on categories with different frequency. We further integrate the proposed Normalized Mask Prediction (Norm Mask) into Mask R-CNN with Seesaw Loss. Without extra cost, the overall AP is improved from 26.6% to 27.1 % and 27.6% to 28.1% with random sampler and RFS, respectively.

Apart from the CE loss baseline, we further compare Seesaw Loss with recent designs for long-tailed instance segmentation, i.e., Equalization Loss (EQL) and Balanced Group Softmax (BAGS) , in Table 1. Seesaw Loss outperforms EQL by 3.9% AP and 1.4% AP, and outperforms BAGS by 1.0% AP and 1.8% AP with random sampler and RFS, respectively. Seesaw Loss also achieves higher APr\text{AP}_{r}, APc\text{AP}_{c} and APf\text{AP}_{f} than the two methods consistently. Notably, EQL and BAGS achieve lower APf\text{AP}_{f} than the CE baseline while Seesaw Loss does not. This phenomenon indicates that these two methods improve the performance of rare and common categories while sacrificing frequent categories.

We further compare Seesaw Loss with previous methods with both random sampler and RFS on Cascade Mask R-CNN . It’s a representative framework of cascade methods that outperforms Mask R-CNN . As shown in Table 1, Seesaw Loss performs much superior to previous works on Cascade Mask R-CNN. Specifically, Seesaw Loss improves the baseline by 6.4% AP and 2.3% AP with random sampler and RFS, respectively. With Normalized Mask Prediction, Cascade Mask-RCNN with Seesaw Loss finally achieves 29.6% AP and 30.1% AP with the two samplers, respectively. Moreover, Seesaw Loss is also evaluated on test-dev split and consistently obtains significant gains over the CE baseline.

3 Ablation study

We conduct a comprehensive ablation study to verify the effectiveness of each design choice in the proposed method.

Components in Seesaw Loss. There are three components in Seesaw Loss: mitigation factor, compensation factor, and normalized linear activation. We evaluated each component on Mask R-CNN with RFS (Table 2). The mitigation factor that mitigates the overwhelming punishments on rare classes leads to a significant improvement from 23.7% AP to 25.1% AP. Notably, it improves the APr\text{AP}_{r} of rare classes by 2.8% AP. The compensation factor increases the punishments of a class when it observes false positives on that class to reduce misclassification. It improves the baseline by 0.4% AP. The combination of the mitigation and the compensation factors achieves 25.7% AP, outperforming the performance of mitigation factor by 0.6% AP. It reveals the effectiveness of instance-wise compensation to avoid misclassification. The normalized linear activation is another important component in Seesaw Loss, which reduces the scale invariance of weights and features across different categories. It improves the baseline performance from 23.7% to 24.7% AP. Seesaw Loss combining all these three components achieves 26.4% AP.

Normalized linear activation. We empirically find that normalized linear activation (NLA) helps to improve the performance of both CE Loss and Seesaw Loss. Therefore, we further integrate NLA with equalization loss (EQL) and balanced group softmax (BAGS) for fair comparisons. Results in Table 3 show that NLA improves the performance of EQL and BAGS by 0.3% and 0.8% AP, respectively. It is noteworthy that Seesaw Loss outperforms EQL and BAGS no matter whether NLA is adopted.

Cumulative Sample Numbers. Different from previous works that rely on the pre-computed frequency distribution of categories in the dataset, Seesaw Loss accumulates the sample numbers of each category during training. We compare different approaches to obtain the sample numbers of categories for Seesaw Loss (Table 4). Directly using the statistics of train split decreases the performance of Seesaw Loss by 0.3% AP. The reason lies in that the data sampler, e.g., repeat factor sampler, changes the frequency distribution of categories during training. We also explore loading the pre-recorded distribution of training samples from a model trained with Seesaw Loss. It achieves a similar performance with accumulating the training samples online (26.3% AP \vs26.4% AP). These results verify the effectiveness and simplicity of online accumulating.

Hyper-parameters. We study the hyper-parameters, i.e., pp, qq, τ\tau, adopted in different components of Seesaw Loss. The normalized linear activation is not applied when studying the mitigation and compensation factors (25.7% AP with this setting). In Table 5, we explore pp in (NjNi)p\left(\frac{N_{j}}{N_{i}}\right)^{p} of mitigation factor. pp controls the magnitude to mitigate the punishments on rare classes. A higher value of pp will reduce punishments more, as well as increase the risk of inducing false positives of tail classes. Therefore, it is critical to find a suitable pp. Results show that p=0.8p=0.8 achieves the best performance. In Table 6, we explore qq in (σjσi)q\left(\frac{\sigma_{j}}{\sigma_{i}}\right)^{q} of the compensation factor. qq controls the magnitude to compensate the reduced punishments on tail classes when false positives are observed. We study the effectiveness of different qq and find q=2.0q=2.0 achieves the best performance. Notably, qq is robust across different values as the best value is only 0.3% AP better than the worst value. In Table 7, we study the temperature τ\tau in normalized linear activation (NLA). τ\tau determines the variance of the classifier’s predicted logit zz. If τ\tau is too small, the variance of zz is insufficient to distinguish positive and negative samples. However, if τ\tau is too big, the target of balancing the variance in weights and features between different categories will be sacrificed. We choose τ=20\tau=20 in the NLA as it achieves the best performance.

Objectness Branch. In a common practice of object detection, the classifier predicts C+1C+1 scores for CC foregrounds categories and one background category. Due to the extremely imbalanced distribution between foregrounds and backgrounds, Seesaw Loss will tend to misclassify more backgrounds as foregrounds with this design. Thus, we adopt an extra objectness branch as described in Sec. 3.3. The results in Table 8 shows that the objectness branch does not improve the Coss-Entropy loss baseline but is critical to Seesaw Loss. The objectness branch helps to avoid reducing the backgrounds’ punishments on CC foreground categories. As a result, the objectness branch brings gains on Seesaw Loss across categories with different frequency, and improves the overall AP from 25.3% to 26.4%.

Training Pipeline. Apart from the end-to-end training pipeline, we further explore the popular decoupling training pipeline on Mask R-CNN. Specifically, we pre-train the Mask R-CNN with Cross-Entropy loss using either random sampler or repeat factor sampler for 2x schedule. Then we finetune the final fully-connected layer of the classifier with all other components fixed. The 1x schedule and repeat factor sampler is adopted during finetuning. As shown in Table 9, Seesaw Loss with the pre-trained model on repeat factor sampler (P-RFS) achieves 25.8% AP, outperforming other methods with decoupling training pipeline. Notably, Seesaw Loss performs better with the end-to-end training pipeline than with the decoupling training pipeline. It indicates Seesaw Loss provides a simpler and more effective solution to long-tailed instance segmentation without relying on complex training pipelines.

4 Long-Tailed Image Classification

To show the versatility of Seesaw Loss, we apply it for long-tailed image classification task on ImageNet-LT dataset. ImageNet-LT is generated from the ImageNet-2012 dataset with long-tailed distributed categories in training set. There are 115.8k images of 1000 categories with a maximum number of 1280 images and a minimum number of 5 images. The performance is evaluated with top-1 accuracy on all categories and the accuracies for Many Shot (>> 100 images), Medium Shot (20∼\sim100 images) and Few Shot (<< 20 images) categories are also reported.

We adopt two training pipelines: end-to-end training and decoupling training. We use ResNeXt-50 backbone and SGD optimizer with momentum of 0.9, initial learning rate of 0.2, batch size of 512, and cosine learning rate following . For the end-to-end training pipeline, the model is trained for 90 epochs. For decouple training pipeline , we load the pre-trained ResNeXt-50 with Cross-Entropy Loss (CE), and finetune the classifier with class-balanced sampler while fixing all other layers for 10 epochs. Seesaw Loss in image classification mostly follows the hyper-parameters on the instance segmentation task except for qq in the compensation factor. We adopt q=1q=1 for ImageNet-LT dataset. The study of qq for ImageNet-LT dataset is shown in Table 11.

We report the performance of Seesaw Loss in Table 10. Seesaw Loss improves top-1 accuracies of CE from 44.4% to 49.7% and 50.4% with the decoupling training and the end-to-end training pipeline, respectively. Similar to our observations on the instance segmentation task, Seesaw Loss performs better with the end-to-end training pipeline on image classification. The performance achieved by Seesaw Loss with the end-to-end pipeline is competitive among previous methods on ImageNet-LT .

Conclusion

In this paper, we propose Seesaw Loss for long-tailed instance segmentation. Seesaw Loss dynamically re-balances gradients of positive and negative samples for each category with two complementary factors. The mitigation factor reduces punishments to tail categories w.r.t. the ratio of cumulative training instances between categories. Meanwhile, the compensation factor increases the penalty of misclassified instances to avoid false positives. Experimental results demonstrate that Seesaw Loss provides a simpler and more effective solution to long-tailed instance segmentation without relying on complex training pipelines.

Acknowledgements. This research was conducted in collaboration with SenseTime. This work is supported by GRF 14203518, ITS/431/18FX, CUHK Agreement TS1712093, NTU NAP, A*STAR through the Industry Alignment Fund - Industry Collaboration Projects Grant, Shanghai Committee of Science and Technology, China (Grant No. 20DZ1100800).

In this work, we propose Seesaw Loss to dynamically re-balance gradients of positive and negative samples for each category. Specifically, Seesaw Loss mitigates the overwhelming gradients of negative samples imposed by a head class ii on a tail class jj via decreasing the value of Sij\mathcal{S}_{ij} in the following formula,

To further analyze the effects of adjusting the value of Sij\mathcal{S}_{ij}, we calculate the partial derivative of Eqn 9 with respect to Sij\mathcal{S}_{ij} as

The value of the partial derivative in Eqn 10 is always positive. This indicates that the gradients of negative samples imposed by class ii on class jj will be reduced as the value of Sij\mathcal{S}_{ij} decreases.

Appendix B How Seesaw Loss works

Via re-balancing gradients of positive and negative samples, Mask R-CNN w/ Seesaw Loss significantly outperforms Mask R-CNN w/ Cross-Entropy Loss on LVIS dataset. Here, we conduct a quantitative analysis of the effectiveness of Seesaw Loss on re-balancing the gradients of positive and negative samples for each category. Specifically, we adopt Mask R-CNN with ResNet-101 backbone and FPN as instance segmentation framework. The Cross-Entropy Loss and Seesaw Loss are integrated into the framework and trained with random sampler by 2x schedule. We accumulate the gradients of positive and negative samples on predicted logit ziz_{i} of each category ii during the whole training procedure.

Figure 3 shows the distribution of the ratio of cumulative gradients between positive and negative samples for each category in Mask R-CNN with Cross-Entropy Loss and Seesaw Loss, respectively. With Cross-Entropy Loss, tail classes obtain heavily imbalanced gradients of positive and negative samples during training. The overwhelming gradients of negative samples lead to a biased learning process for the classifier, which results in the low classification accuracy on tail classes. On the contrary, Seesaw Loss effectively re-balances the gradients of positive and negative samples across different categories. Consequently, Mask R-CNN with Seesaw Loss achieves significant improvements on instance segmentation performance as shown in Figure 1 and Table 1 in the main text.

Appendix C Per-category Performance Comparison

In addition to the performance reported in Table 1 of the main text, we further show the per-category performance (AP) to verify the superiority of Seesaw Loss compared to other loss functions. As shown in Figure 4, compared to other loss functions (i.e., Cross-Entropy Loss, Equalization Loss , and Balanced Group Softmax ), Seesaw Loss consistently achieves strong performance across categories with different frequency on different frameworks (i.e. Mask R-CNN , Cascade Mask R-CNN ) and samplers (i.e., random sampler, repeat factor sampler ).

Appendix D LVIS Challenge 2020

Here we present the approach used in the entry of team MMDet in the LVIS Challenge 2020. In our entry, we adopt Seesaw Loss for long-tailed instance segmentation as described in the main text. Seesaw Loss improves the strong baseline by 6.9% AP on LVIS v1 val split. Furthermore, we propose HTC-Lite, a light-weight version of Hybrid Task Cascade (HTC) which replaces the semantic segmentation branch with a global context encoder. With a single model and without using external data and annotations except for standard ImageNet-1k classification dataset for backbone pre-training, our entry achieves 38.92% AP on the test-dev split of the LVIS v1 benchmark.

We propose HTC-Lite, a light-weight version of Hybrid Task Cascade (HTC) , to accelerate the training and inference speed while maintaining good performance. As shown in Figure 5, the modifications are in two folds: replacing the semantic segmentation branch with a global context encoding branch and reducing mask heads.

Context Encoding Branch. Since semantic segmentation annotations are not available for LVIS dataset, we replace the semantic segmentation branch with a global context encoder which works as a multi-label classification branch trained by a binary cross-entropy loss. The context encoder applies convolution layers and a global average pooling on the input feature map to obtain a feature vector. And an auxiliary fully connected (fc) layer is applied on the feature vector to predict the categories existing in the current image. By this approach, this feature vector encodes the global context information of the image. Then it is added to the RoI features used by box heads and mask heads to enrich their semantic information.

Reduced Mask Heads. To further reduce the cost of instance segmentation, HTC-Lite only keeps the mask head in the last stage, which also spares the original interleaved information passing.

In Table 13, we compare the performance and inference speed on LVIS v1 dataset of HTC-Lite with two mainstream cascading instance segmentation frameworks, i.e., Cascade Mask R-CNN and HTC. The ResNet-50 with FPN backbone, repeat factor sampler and 1x training schedule are adopted in these methods. The semantic segmentation branch in HTC is removed since semantic segmentation annotations are not available on LVIS v1 dataset. We evaluate the inference speed for each framework with a single Tesla V100 GPU. The experimental results show that HTC-Lite is not only much more efficient than its counterparts but also outperforms them.

D.2 Step by Step Results

We report the step-by-step results of our entry in LVIS Challenge 2020 as shown in Table 12.

Baseline. The baseline model is Mask R-CNN using ResNet-50-FPN , trained with multi-scale training and random data sampler by 2x schedule .

SyncBN. We use SyncBN in the backbone and heads.

CARAFE Upsample. CARAFE is used for upsampling in the mask head.

HTC-Lite. We use HTC-Lite as described in Appendix D.1.

TSD. TSD is used to replace the box heads in all three stages in HTC-Lite.

Mask Scoring. We further use the mask IoU head to improve mask results.

Training Time Augmentation. We train the model with stronger augmentations with 45 epochs. The learning rate is decreased by 0.1 at 30 and 40 epochs. We randomly resize the image with its longer edge in a range of 768 to 1792 pixels. And then, we randomly crop the image to the size of 1280×12801280\times 1280 after adopting instaboost augmentation .

Stronger Neck. We replace the neck architecture with an enhanced version of Feature Pyramid Grids (FPG) . The enhanced FPG uses deformable convolution v2 (DCNv2) after feature upsampling, and a downsampler version of CARAFE for feature downsampling.

Stronger Backbone. We use ResNeSt-200 with DCNv2 as the backbone.

Seesaw Loss. We apply the proposed Seesaw Loss to classification branches of the TSD box head, in all cascading stages. Furthermore, we remove the original progressive constraint (PC) loss on classification branches in TSD.

Dual Head Classification. Inspired by , we adopt a dual-head classification policy to further boost the performance. Specifically, after obtaining the model with Seesaw Loss trained by a random sampler, we freeze all components in the original model. Then we finetune a new classification branch for each cascading stage on the fixed model using repeat factor sampler by 1x schedule. During inference, the classification scores of original classification branches and the scores of new classification branches are averaged to get the final scores.

Test Time Augmentation. We adopt multi-scale testing with horizontal flipping. Specifically, image scales are 1200, 1400, 1600, 1800, and 2000 pixels.

Final Performance on Test-dev. After adding the abovementioned components step by step, we finally achieve 38.8% AP on the val split and 38.92% AP on the test-dev split.

References