Exploring Classification Equilibrium in Long-Tailed Object Detection
Chengjian Feng, Yujie Zhong, Weilin Huang
Introduction
Object detection plays an important role in computer vision, and recent object detectors have achieved promising performance on several common datasets with a few categories and balanced class distribution, such as PASCAL VOC (20 classes) and COCO (80 classes) . However, most real-world data contains a large number of categories and its distribution is long-tailed: a few head classes contain abundant instances while a great number of tail classes only have a few instances.
Recently, LVIS is released for exploring long-tailed object detection. Not surprisingly, the performance of the state-of-the-art detectors designed for balanced data is significantly degraded if they are directly applied to such datasets. The reason for the performance degradation mainly comes from two aspects: (1) The long-tailed distribution of data. The number of the instances from tail classes (e.g., only 1 instance for one class) is insufficient for training a deep learning model, resulting in under-fitting of these classes. Moreover, the tail classes will be overwhelmed by the head classes during training, because the number of the instances of head classes is much larger than that of tail classes (e.g., thousands of times). As a result, the detectors cannot learn the tail classes well, and recognize those tail classes with very low confidence, as demonstrated in Figure 1. (2) The large number of categories. With the increase of the number of categories, it brings a higher chance of misclassification, especially for the tail classes with a very low classification score.
Several works attempted to cope with the problem of long-tail learning by re-sampling training data or re-weighting loss function. However, most of them assign the sampling rate and the loss weight according to the sampling frequency of each category, which is model-agnostic and sensitive to hyper-parameters . It may bring the following problems: (1) the model-agnostic data re-sampling is prone to over-fit the tail classes and under-represent the head classes; (2) the dataset-based loss re-weighting may cause excessive gradients and unstable training especially when the category distribution is extremely imbalanced. Recently, wang et al. introduce Seesaw loss to adaptively re-balance the gradients of positive and negative samples by dynamically accumulating the number of class instances during training. However, the number of the training samples cannot accurately reflect the learning quality of the classes, due to the diversity and complexity of instances and categories, e.g., training a classifier for the categories with visual similarity usually requires more training samples than the categories with very different visual appearance.
To address the above problems, we propose to use the mean classification score to monitor the learning status (i.e., classification accuracy) of each category during training. As shown in Figure 1, the mean classification score has an approximate positive correlation with the classification accuracy. Thus, it can be used as an effective indicator to reflect the classification accuracy during training. Based on this indicator, we design an Equilibrium Loss (EBL) and a Memory-augmented Feature Sampling (MFS) method, to dynamically balance the classification. Equilibrium Loss: To balance the classification of different classes, EBL assigns different loss margins between any two classes based on the statistical mean classification score. It increases the loss margin between weak (with low mean score) positive classes and dominant (with high mean score) negative classes, and vice versa. Thus, the designed loss margin increases the intensity of the adjustment of the classification decision boundary for the weak classes, resulting in a more balanced classification. Memory-augmented Feature Sampling: In addition to increasing the intensity of the adjustment of the decision boundary, we design MFS to increase the frequency and accuracy of the adjustment of the decision boundary for the weak classes. Specifically, rich instance features are firstly extracted based on a set of dense bounding boxes generated by a model-agnostic bounding box generator, and then stored by a feature memory module for feature reuse across training iterations. Finally, a probabilistic sampler is used to access the feature memory module to sample more instance features of weak classes to improve the training.
We codename the proposed method Long-tailed Object detector with Classification Equilibrium (LOCE). In summary, our contributions are as follows: (1) we propose to use the mean classification score to monitor the classification accuracy of each category during training; (2) we develop a score-guided equilibrium loss that improves the intensity of the adjustment of the decision boundary for the weak classes. (3) we design a memory-augmented feature sampling to enhance the frequency and accuracy of the adjustment of the decision boundary for the weak classes. (4) we conduct experiments on LVIS using Mask R-CNN with various backbones including ResNet-50-FPN and ResNet-101-FPN . Extensive experiments show the superiority of LOCE. It improves the tail classes by 15.6 AP based on the Mask R-CNN with ResNet-50-FPN and outperforms the most recent long-tailed object detectors by more than 1 AP on LVIS v1.0.
Related Work
Modern object detection frameworks can be divided into two-stage and one-stage ones. The two-stage detectors first generate a set of region proposals, and then classify and refine the proposals. In contrast, the one-stage detectors directly predict the category and bounding box at each location. Most of the detectors are designed for balanced data. When it comes to long-tailed data, the performance of such detectors is greatly degraded. Recently, extensive studies attempted to optimize the two-stage detectors such as to cope with the long-tailed data, by designing balanced samplers or balanced loss functions . Some works adopt decoupled training pipeline , which first learns the universal representations with unbalanced data and then fine-tunes the classifier with re-balanced data or balanced loss function. Inspired by them, we propose both adaptive feature sampling and adaptive loss function for long-tailed detection.
Sampler for long-tail learning.
Data re-sampling is a common solution for long-tail learning. It typically over-samples the training data from tail classes while under-samples those from head classes. In long-tailed detection, the data samplers balance the training data on the image-level or instance-level. Gupta et al. use image-level Repeat Factor Sampling (RFS) to up-sample the data from minority classes based on the sampling frequency of each class. Wang et al. propose a class-based sampler to balance the data from the instance-level, by only considering the proposals of the selected classes. Wu et al. set a higher NMS threshold for tail classes to sample more proposals from tail classes. These methods design the balanced sampler depending on the frequency distribution of categories. In contrast, we design a memory-augmented feature sampling based on the mean classification score, which can adapt to the training process dynamically. Recently, ren et al. introduce Meta Sampler to estimate the optimal sample rate with meta-learning. Compared with Meta Sampler, our feature sampling is simpler and more versatile.
Loss Function for long-tail learning.
Balanced loss functions have received lots of attention in long-tailed classification. Most of them are achieved through loss weighting or margin modification related to the distribution of training data. For example, the works such as re-weight the loss functions by the inverse of the sampling frequency of each class, while those of increase the loss margins of tail classes and decrease those of head classes for balanced classification. Recently, several works attempted to design balanced loss for long-tailed detection. Tan et al. propose Equalization Loss (EQL) to improve the performance of tail classes by ignoring the suppressing gradients for tail classes. Li et al. introduce Balanced Group Softmax that first groups the classes based on the instance numbers and then separately apply softmax within each group. Ren et al. design Balanced Softmax to accommodate the label distribution shifts of the long-tailed data according to the number of category samples. Tan et al. improve EQL by re-balancing the positive and negative gradients for each category independently and equally. Wang et al. develop Seesaw loss to re-balance the gradients of positive and negative samples by accumulating the number of the training samples. Different from the existing methods, we design the loss function according to the mean classification score calculated during training. It can track and adjust the learning status of the model dynamically.
Methodology
As mentioned in Section 1, if the distribution of the training data is severely skewed, the mean classification score obtained by the conventional detectors is extremely imbalanced for each category. In this work, we propose LOCE, an object detector with classification equilibrium, to alleviate this problem. We first use the mean classification score to indicate the learning status (i.e., classification accuracy) of each category during training (Section 3.1). Then, based on this indicator, we balance the classification through a score-guided equilibrium loss (Section 3.2) and a memory-augmented feature sampling method (Section 3.3). The proposed loss function and the feature sampling method collaboratively adjust the classification decision boundary, as demonstrated in Figure 2. Similar to most methods for long-tailed object classification or detection, we adopt the decoupled training pipeline . The two methods are adopted in the fine-tuning stage.
We first analyze the classification problem when applying conventional detectors to long-tailed data, and then introduce the mean classification score to indicate the learning status of the detector.
Classification accuracy indicator.
To alleviate the problem of classification imbalance, we attempt to find an effective indicator to reflect the learning status (i.e., classification accuracy) of the classifier for each category and dynamically adjust the learning process. Previous works proposed to balance the classification based on the number of training samples of each category. However, the number of the training samples cannot indicate the learning quality of the model accurately, because of the diversity and complexity of instances and categories. For example, training a classifier for the categories with high inter-class visual similarity usually requires more training samples than the categories with very different visual appearances. Instead, we seek a more effective indicator to reflect the classification status. From the statistics shown in Figure 1, we found that the mean classification score has an approximate positive correlation with the classification performance. Namely, for LVIS, the head classes have higher mean classification scores and higher classification accuracy, while the tail classes have lower mean classification scores and lower classification accuracy. For the balanced dataset COCO, we also observe a similar pattern: high mean classification scores are usually associated with high classification accuracy.
where is the predicted probability of the instance, and is a smoothing coefficient hyper-parameter.
Compared with the existing works that use the number of training instances to indicate the learning status of the model for each category, the proposed indicator has the following advantages: (1) it can monitor the classification accuracy of each category during training; (2) it can be applied when the distribution of the training data is not visible or the model is pre-trained with other datasets, e.g., the training samples are obtained from an online stream. The mean classification score affects the training by guiding the proposed loss function and the feature sampling method to balance the classifier, which is introduced in the following.
2 Equilibrium Loss
where works as a tunable balancing margin between any two classes, based on the distribution of the mean classification score. To adjust the decision boundary accordingly, the design of should satisfy the following two properties: (1) it should reduce the suppression of dominant classes (i.e., having high mean classification score) over weak classes (i.e., having low mean classification score), which can be achieved by reducing the margin between dominant positive classes and weak negative classes; (2) it should enlarge the suppression of weak classes over dominant classes, which can be achieved by increasing the margin between weak positive classes and dominant negative classes. Therefore, we design the following adaptive loss margin between any two classes:
As defined above, the background is regarded as an auxiliary category in the classifier. In the experiments, we found that the classifier trained with EBL tends to predict false positive results, i.e., misclassifying the backgrounds as foregrounds. To reduce those false positive cases, we enlarge the corresponding punishment. In the margin loss, it can be achieved by increasing the margin between the positive background class and the negative foreground classes. Consequently, we decrease the mean classification score of background class when computing Eq.(4). It is worth noting that most training samples (at least 75%) for classifier are negative samples (i.e., background). For simplicity and efficiency, instead of computing the statistical , we use a small value (e.g., 0.01) to replace , namely . Some works such as attempt to reduce the false positive cases by introducing an extra objectness branch. Comparatively, the proposed method is simpler and more efficient.
3 Memory-augmented Feature Sampling
Although EBL tends to move the decision boundary from the tail classes to the dominant ones, the decision boundary sometimes is still closer to the tail classes. This is because EBL only adjusts the intensity of the adjustment of the decision boundary, while the frequency of the adjustment is very low for tail classes during training, due to the overwhelming number of images in the head classes. Especially, the tail classes usually have very few training samples (e.g., for LVIS). Therefore, we need to increase the number (or occurrence) of the training samples of tail classes to enhance the frequency of such boundary adjustment, in addition to EBL.
To increase the occurrence of the training samples from tail classes, a straightforward method is data re-sampling. As shown in Figure 3(a), the widely used sampling methods for data balancing can be divided into two categories: (1) Dataset-based image sampling such as Repeat Factor Sampling (RFS) and Class-balanced Sampling (CBS) usually samples more training images of the tail classes based on the training set statistics. (2) Class-based proposal sampling such as NMS Resampling and Bi-level Sampling samples more proposals from Region Proposal Network (RPN) for the tail classes or the selected classes. Despite their success, these sampling methods have the following limitations: (1) the diversity of the training features from proposal sampling depends on the behavior of RPN; (2) image sampling usually requires more training iterations; (3) most existing sampling methods are model-agnostic and prone to over-fit the tail classes and under-represent the head classes. To overcome these limitations, we propose a more efficient Memory-augmented Feature Sampling (MFS) method. As shown in Figure 3(b), MFS consists of a bounding box generator, a feature memory module, and a probabilistic sampler.
Recent two-stage detectors sample RoI features for the classifier based on the proposals from RPN and the sampling configuration. In such pipeline, the diversity of the classification training features is limited by the performance of RPN and the sampling configuration such as the number of positive samples at each iteration. It hinders the classifier from improving its generalization ability and accuracy, especially for the tail classes. Instead, we design a model-agnostic bounding box generator with aim to extract rich instance features for training the classifier. Moreover, we propose to reuse the extracted instance features across training iterations. With such design, the feature diversity for the classifier is no longer sensitive to the performance of RPN and the sampling configuration.
Concretely, given an object instance, we have its ground-truth class and ground-truth box , where and are the coordinates of the upper left corner and the lower right corner of the ground-truth box. The bounding box generator yields the following dense bounding boxes:
where , are box width and height, i.e., , , and is a random number. Such bounding box generator can obtain any potential positive bounding boxes (i.e., IoU 0.5). Based on the dense bounding boxes , we extract the corresponding instance features by applying RoI-Align to the features from Feature Pyramid Network (FPN).
Feature Memory Module.
Image re-sampling usually requires extra training iterations while proposal sampling does not provide enough samples to balance the classification, especially for tail classes. In contrast, the proposed bounding box generator can yield dense bounding boxes for extracting more instance features, without the requirement of extra training iterations. However, using all the instance features from the dense bounding boxes to train the classifier within an iteration can bring lots of computational overhead and memory consumption, especially when there are lots of instances in an image. Besides, for the tail classes that only occur a few times in a training epoch, taking all the instance features within an iteration may still be inadequate.
To alleviate the above problems, we use a feature memory module to store the instance features, and reuse the instance features as needed during the following training. The rationale is that the parameters of the backbone (e.g., ResNet-50-FPN) are frozen during the fine-tuning stage. Thus, the features extracted from the backbone are stable and can be reused to fine-tune the classifier during different training iterations. Memory module is widely adopted in contrastive learning , but never used in long-tailed object detection. Specifically, the feature memory module is maintained and updated with a class queue for each class :
Probabilistic Sampler.
At each training iteration, we access the feature memory module by a sampler to augment the training features for balancing the classifier. In particular, we use the mean classification score to adaptively adjust the sampling process, similar to that for EBL. To improve the training effectiveness and classification score of the weak classes (e.g., tail classes), we sample more memory features of these classes. Specifically, we design a probabilistic sampler that samples the memory features according to the probability with negative correlation with the mean classification score, namely:
where is a non-increasing transform and is the sampling probability of class . For simplicity, we define:
Finally, we randomly choose classes according to , and select features from the feature memory for each selected class. Then, the selected features will be used together with the RoI features from RPN to train the classifier.
Experiments
We conduct experiments on the recent long-tail and large-scale dataset LVIS . The latest version v1.0 contains 1203 categories with both bounding box and instance mask annotations. We use the set (100k images with 1.3M instances) for training and set (19.8k images) for validation. We also perform experiments on LVIS v0.5, which contains 1230 and 830 categories respectively in its set and set. All the categories are divided into three groups based on the number of the images that each category appears in the set: rare (1-10 images), common (11-100 images), and frequent (100 images). Apart from the official metrics Average Precision (AP), we also report APr (for rare classes), APc (for common classes) and APf (for frequent classes) to measure the detection performance and segmentation performance. Unless specific, APb denotes the detection performance, while AP denotes the segmentation performance.
Implementation details.
We implement our method with MMDetection and conduct experiments using Mask R-CNN with various backbones including ResNet-50-FPN and ResNet-101-FPN pre-trained on ImageNet . Following , we use the decoupled training pipeline. Namely, we first train the model with standard softmax cross-entropy and random image sampler for 24 epochs, then fine-tune the model with the proposed method for 6 epochs. Specifically, the initial learning rate is 0.02 and dropped by a factor of 10 at the th and th epoch for the first training stage and the th and th epoch for the fine-tuning stage. The models are trained using SGD optimizer with 0.9 momentum and 0.0001 weight decay and batch size of 16 on 8 GPUs. Following the convention, we train the detector with scale jitter (640-800) and horizontal flipping. At testing time, the model is evaluated without test time augmentation, and the maximum number of detections per image is 300 with the minimum score threshold of 0.0001. We set as 0.9 for updating the mean classification score. The substituted is set to be 0.01. The memory size is 80, and and in feature sampler are 8 and 4 on each GPU. As in , we adopt normalized linear activation to the mask prediction. More implementation and training details refer to the supplementary material.
2 Ablation Study
We use Mask R-CNN with backbone ResNet-50-FPN for ablation study, and report the results on LVIS v1.0.
Table 1 reports the detection and segmentation results of each proposed component. For a fair comparison, we train the baseline using standard softmax cross-entropy and random image sampler for 30 epochs. First, we evaluate the performance of EBL. EBL improves both the APb and AP by 3.6 AP, comparing to the baseline. Specifically, it improves the performance of the classes of all groups, i.e., +5.4 AP for rare classes, +5.5 AP for common classes and +0.6 AP for frequent classes, respectively. These results show that the score-guided loss margin can also help to optimize the decision boundary of head classes, even though it is mainly designed for weak classes. We then examine the effectiveness of MFS. Compared with the baseline, MFS improves the performance by 5.6 AP for object detection and 5.3 AP for instance segmentation. To be more specific, most of the improvements are from rare classes and common classes, which yield +15.5 AP and +6.8 AP improvement for instance segmentation. We can see that using more instance features from weak classes can bring large improvements, especially for rare classes. Next, we verify the effectiveness of the complete method (i.e., LOCE). EBL and MFS work collaboratively and improve the AP by 7.0 AP for object detection and 6.4 AP for instance segmentation, comparing to the baseline. Notably, MFS itself achieves a little lower performance on frequent classes than the baseline while LOCE achieves higher performance on frequent classes. This shows that EBL helps MFS to find a better equilibrium point for the frequent classes. Therefore, LOCE dramatically improves the performance of tail classes, while maintaining or even improving the performance of head classes.
Hyper-parameters.
We compare the performance with different smoothing coefficients for updating the mean classification score. From the experiment results shown in Table 2, we observe that the performance is insensitive to the value of , and yields the best performance. The performance with different is shown in Table 3. We can see that the detector achieves the best performance when . We then conduct several experiments to study the robustness with respect to and of feature sampler in Table 4. Through a coarse search, we set and for the rest of the experiments.
Analysis of classification equilibrium.
Here we analyze the mean classification score and the classification accuracy between different methods on LVIS set. As shown in Figure 4, the distribution of the mean classification score predicted by the detector trained with softmax cross-entropy loss and random image sampling is severely skewed. Specifically, the mean classification score of tail classes is close to 0, and its classification accuracy is also close to 0. When using RFS instead of random image sampling to train the detector, both the mean classification score and the classification accuracy are improved marginally. In contrast, the mean classification score predicted by the proposed LOCE is more balanced than those predicted by the two methods mentioned above, and the classification accuracy of common classes and tail classes is improved.
3 Comparison with the State-of-the-Art
We compare LOCE with the state-of-the-art methods on LVIS v0.5 and LVIS v1.0 in Table 5. On LVIS v0.5, the proposed method achieves the detection performance of 28.2 AP and segmentation performance of 28.4 AP, surpassing the most recent long-tailed object detectors such as BAGS (by 2.4 AP and 2.1 AP) and BALMS (by 0.6 AP and 1.4 AP). Specifically, it outperforms BAGS by 4.0 points on rare classes, which shows the superior performance of the proposed method for tail classes. Compared with most existing methods such as , the proposed method gets a much higher performance on head classes, in addition to improving the detection performance on tail classes. On LVIS v1.0, the proposed method achieves a better result than all the methods shown in Table 5, including the concurrent work such as Seesaw Loss and EQL v2 . With the framework of Mask R-CNN , LOCE achieves the detection performance of 27.4 AP and 29.0 AP on R-50-FPN and R-101-FPN, outperforming recent works such as BAGS , Seesaw Loss and EQL v2 by more than 1 AP.
Conclusion
In this paper, we explore the classification equilibrium in long-tailed object detection. We propose to use the mean classification score to indicate the learning status of the model for each category, and design an equilibrium loss and a memory-augmented feature sampling method to balance the classification. Extensive experiments show the superiority of the proposed method, which sets a new state-of-the-art in long-tailed object detection.