On Model Calibration for Long-Tailed Object Detection and Instance Segmentation

Tai-Yu Pan, Cheng Zhang, Yandong Li, Hexiang Hu, Dong Xuan, Soravit Changpinyo, Boqing Gong, Wei-Lun Chao

Introduction

Object detection and instance segmentation are the fundamental tasks in computer vision and have been approached from various perspectives over the past few decades . With the recent advances in neural networks , we have witnessed an unprecedented breakthrough in detecting and segmenting frequently seen objects such as people, cars, and TVs . Yet, when it comes to detect rare, less commonly seen objects (e.g., walruses, pitchforks, seaplanes, etc.) , there is a drastic performance drop largely due to insufficient training samples . How to overcome the “long-tailed” distribution of different object classes has therefore attracted increasing attention lately .

To date, most existing works tackle this problem in the model training phase, e.g., by developing algorithms, objectives, or model architectures to tackle the long-tailed distribution . Wang et al. investigated the widely used instance segmentation model Mask R-CNN and found that the performance drop comes primarily from mis-classification of object proposals. Concretely, the model tends to give frequent classes higher confidence scores , hence biasing the label assignment towards frequent classes. This observation suggests that techniques of class-imbalanced learning can be applied to long-tailed detection and segmentation.

Building upon the aforementioned observation, we take another route in the model inference phase by explicit post-processing calibration , which adjusts a classifier’s confidence scores among classes, without changing its internal weights or architectures. Post-processing calibration is efficient and widely applicable since it requires no re-training of the classifier. Its effectiveness on multiple imbalanced classification benchmarks may also translate to long-tailed object detection and instance segmentation.

In this paper, we propose a simple post-processing calibration technique inspired by class-imbalanced learning and show that it can significantly improve a pre-trained object detector’s performance on detecting both rare and common classes of objects. We note that our results are in sharp contrast to a couple of previous attempts on exploring post-processing calibration in object detection , which reported poor performance and/or sensitivity to hyper-parameter tuning. We also note that the calibration techniques in are implemented in the training phase and are not post-processing.

Concretely, we apply post-processing calibration to the classification sub-network of a pre-trained object detector. Taking Faster R-CNN and Mask R-CNN for examples, they apply to each object proposal a (C+1)(C+1)-way softmax classifier, where CC is the number of foreground classes, and 1 is the background class. To prevent the scores from being biased toward frequent classes , we re-scale the logit of every class according to its class size, e.g., number of training images. Importantly, we leave the logit of the background class intact because (a) the background class has a drastically different meaning from object classes and (b) its value does not affect the ranking among different foreground classes. After adjusting the logits, we then re-compute the confidence scores (with normalization across all classes, including the background) to decide the label assignment for each object proposalPopular evaluation protocols allow multiple labels per proposal if their confidence scores are high enough. (see Figure 1). We note that it is crucial to normalize the scores across all classes since it triggers re-ranking of the detection results within each class (see Figure 3), influencing the class-wise precision and recall. Instead of separately adjusting each class by a specific factor , we follow to set the factor as a function of the class size, leaving only one hyper-parameter to tune. We find that it is robust to use the training set to set this hyper-parameter, making our approach applicable to scenarios where collecting a held-out representative validation set is challenging.

Our approach, named Normalized Calibration for long-tailed object detection and instance segmentation (NorCal), is model-agnostic as long as the detector has a softmax classifier or multiple binary sigmoid classifiers for the objects and the background. We validate NorCal on the LVIS dataset for both long-tailed object detection and instance segmentation. NorCal can consistently improve not only baseline models (e.g., Faster R-CNN or Mask R-CNN ) but also many models that are dedicated to the long-tailed distribution. Hence, our best results notably advance the state of the art. Moreover, NorCal can improve both the standard average precision (AP) and the category-independent APFixed{}^{\text{Fixed}} metric , implying that NorCal does not trade frequent class predictions for rare classes but rather improve the proposal ranking within each class. Indeed, through a detailed analysis, we show that NorCal can in general improve both the precision and recall for each class, making it appealing to almost any existing evaluation metrics. Overall, we view NorCal a simple plug-and-play component to improve object detectors’ performance during inference.

Related Work

Long-tailed detection and segmentation. Existing works on long-tailed object detection can roughly be categorized into re-sampling, cost-sensitive learning, and data augmentation. Re-sampling methods change the long-tailed training distribution into a more balanced one by sampling data from rare classes more often . Cost-sensitive learning aims at adjusting the loss of data instances according to their labels . Building upon these, some methods perform two- or multi-staged training , which first pre-train the models in a conventional way, using data from all or just the head classes; the models are then fine-tuned on the entire long-tailed data using either re-sampling or cost-sensitive learning. Besides, another thread of works leverages data augmentation for the object instances of the tail classes to improve long-tailed object detection .

In contrast to all these previous works, we investigate post-processing calibration to adjust the learned model in the testing phase, without modifying the training phase or modeling. Concretely, these methods adjust the predicted confident scores (i.e., the posterior over classes) for each test instance, e.g., by normalizing the classifier norms or by scaling or reducing the logits according to class sizes . Post-processing calibration is quite popular in imbalanced classification but not in long-tailed object detection. To our knowledge, only Li et al. and Tang et al. have studied this approach for object detectionCalibration in is in the training phase and is not post-processing. . Li et al. applied classifier normalization as a baseline but showed inferior results; Tang et al. developed causal inference calibration rules, which however require a corresponding de-confounded training step. Dave et al. applied methods for calibrating model uncertainty, which are quite different from class-imbalanced learning (see the next paragraph). In this paper, we demonstrate that existing calibration rules for class-imbalanced learning can significantly improve long-tailed object detection, if paired with appropriate ways to deal with the background class and normalized the adjusted logits. We refer the reader to the supplementary material for a comprehensive survey and comparison of the literature.

Calibration of model uncertainty. The calibration techniques we employ are different from the ones used for calibrating model uncertainty : we aim to adjust the prediction across classes, while the latter adjusts the predicted probability to reflect the true correctness likelihood. Specifically for long-tailed object detection, Dave et al. applied techniques for calibrating model uncertainty to each object class individually. Namely, a temperature factor or a set of binning grids (i.e., hyper-parameters) has to be estimated for each of the hundreds of classes in the LVIS dataset, leaving the techniques sensitive to hyper-parameter tuning. Indeed, Dave et al. showed that it is quite challenging to estimate those hyper-parameters for tail classes. In contrast, the techniques we apply have only a single hyper-parameter, which can be selected robustly from the training data.

Post-Processing Calibration for Long-Tailed Object Detection

In this section, we provide the background and notation for long-tailed object detection and instance segmentation, describe our approach Normalized Calibration (NorCal), and discuss its relation to existing post-processing calibration methods.

Our tasks of interests are object detection and instance segmentation. Object detection focuses on detecting objects via bounding boxes while instance segmentation additionally requires precisely segmenting each object instance in an image. Both tasks involve classifying the object in each box/mask proposal region into one of the pre-defined classes. This classification component is what our proposed approach aims to improve. The most common object classification loss is the cross-entropy (CE) loss,

where y∈{0,1}C+1\bm{y}\in\{0,1\}^{C+1} is the one-hot vector of the ground-truth class and p(c∣x)p(c|{\bm{x}}) is the predicted probability (i.e., confidence score) of the proposal x{\bm{x}} belonging to the class cc, which is of the form

Here, ϕc\phi_{c} is the logit for class cc, which is usually realized by wc⊤fθ(x)\bm{w}_{c}^{\top}f_{\bm{\theta}}({\bm{x}}): wc\bm{w}_{c} is the linear classifier associated with class cc and fθf_{\bm{\theta}} is the feature network. We use C+1C+1 to denote the “background” class.

During testing, a set of “(box/mask proposal, object class, confidence score)” tuples are generated for each image; each proposal can be paired with multiple classes and appears in multiple tuples if the corresponding scores are high enough. The most common evaluation metric for these tuples is average precision (AP), where they are compared against the ground-truths for each classThe difference between AP for object detection and instance segmentation lies in the computation of the intersection over union (IoU): the former based on boxes and the latter based on masks.. Concretely, the tuples with predicted class cc will be gathered, sorted by their scores, and compared with the ground-truths for class cc. Further, for popular benchmarks such as MSCOCO and LVIS , there is a cap KK (often set to 300300) on the number of detected objects per image, which is enforced usually by including only the tuples with top KK confidence scores. Such a cap makes sense in practice, since a scene seldom contains over 300300 objects; creating too many, likely noisy tuples can also be annoying to users (e.g., for a camera equipped with object detection).

Long-tailed object detection and instance segmentation: problems and empirical evidence. Let NcN_{c} denote the number of training images of class cc. A major challenge in long-tailed object detection is that NcN_{c} is imbalanced across classes, and the learned classifier using Eq. 1 is biased toward giving higher scores to the head classes (whose NcN_{c} is larger) . For instance, in the long-tailed object detection benchmark LVIS whose classes are divided into frequent (Nc>100N_{c}>100), common (100≥Nc>10100\geq N_{c}>10), and rare (Nc≤10N_{c}\leq 10), the confidence scores of the rare classes are much smaller than the frequent classes during inference (see Figure 2). As a result, the top KK tuples mostly belong to the frequent classes; proposals of the rare classes are often mis-classified as frequent classes, which aligns with the observations by Wang et al. .

2 Normalized Calibration for Long-tailed Object Detection (NorCal)

Next, we describe the key components of the proposed NorCal, including confidence score calibration and normalization. The former re-ranks the confidence scores across classes to overcome the bias that rare classes usually have lower confidence scores; the latter helps to re-order the scores of detected tuples within each class for further improving the performance. Both the confidence score calibration and normalization are essential to the success of NorCal.

Post-processing calibration and foreground-background decomposition. We explore applying simple post-calibration techniques from standard multi-way classification to object detection and instance segmentation. The main idea is to scale down the logit of each class cc by its size NcN_{c} . In our case, however, the background class poses a unique challenge. First, NC+1N_{C+1} is ill-defined since nearly all images contain backgrounds. Second, the background regions extracted during model training are drastically different from the foreground object proposals in terms of amounts and appearances. We thus propose to decompose Eq. 2 as follows,

where the first term on the right-hand side predicts how likely x{\bm{x}} is foreground (vs. background, i.e., class C+1C+1) and the second term predicts how likely x{\bm{x}} belongs to class cc given that it is foreground. Note that, the background logit ϕC+1(x)\phi_{C+1}({\bm{x}}) only appears in the first term and is compared to all the foreground classes as a whole. In other words, scaling or reducing it does not change the order of confidence scores among the object classes c∈{1,⋯ ,C}c\in\{1,\cdots,C\}. We thus choose to keep ϕC+1(x)\phi_{C+1}({\bm{x}}) intact. Please refer to Section 4 for a detailed analysis, including the effect of adjusting ϕC+1(x)\phi_{C+1}({\bm{x}}).

For the foreground object classes, inspired by Figure 2 and the studies in , we propose to scale down the exponential of the logit ϕc(x),∀c∈{1,⋯ ,C}\phi_{c}({\bm{x}}),\forall c\in\{1,\cdots,C\}, by a positive factor aca_{c},

in which aca_{c} should monotonically increase with respect to NcN_{c} — such that the scores for head classes will be suppressed. We investigate a simple way to set aca_{c}, inspired by ,

which has a single hyper-parameter γ\gamma that controls the strength of dependency between aca_{c} and NcN_{c}. Specifically, if γ=0\gamma=0, we recover the original confidence scores in Eq. 2. We investigate other methods beyond Eq. 4 and Eq. 5 in Section 4.

Hyper-parameter tuning. Our approach only has a single hyper-parameter γ\gamma to tune. We observe that we can tune γ\gamma directly on the training dataUnlike imbalanced classification in which the learned classifier ultimately achieves ∼100%\sim 100\% accuracy on the training data (so hyper-parameter tuning using the training data becomes infeasible), a long-tailed object detector can hardly achieve 100%100\% AP per class even on the training data., bypassing the need of a held-out set which can be hard to collect due to the scarcity of examples for the tail classes. Dave et al. also investigate this idea; however, the selected hyper-parameters from training data hurt the test results of rare classes. We attribute this to the fact that their methods have separate hyper-parameters for each class, and that makes them hard to tune.

The importance of normalization and its effect on AP. At first glance, our approach NorCal seems to simply scale down the scores for head classes, and may unavoidably hurt their AP due to the decrease of detected tuples (hence the recall) within the cap. However, we point out that the normalization operation (i.e., sum to 1) in Eq. 4 can indeed improve AP for head classes — normalization enables re-ordering the scores of tuples within each class.

Let us consider a three-class example (see Figure 3), in which c=1c=1 is a tail class, c=2c=2 and c=3c=3 are head classes, and c=4c=4 is the background class. Suppose two proposals are found from an image: proposal AA has scores [0.0,0.4,0.5,0.1][0.0,0.4,0.5,0.1] and the true label cGT=3c_{GT}=3; proposal BB has scores [0.3,0.0,0.6,0.1][0.3,0.0,0.6,0.1] and the true label cGT=1c_{GT}=1. Before calibration, proposal BB is ranked higher than AA for c=3c=3, resulting in a low AP. Let us assume a1=1a_{1}=1 and a2=a3=4a_{2}=a_{3}=4. If we simply divide the scores of object classes by these factors, proposal BB will still be ranked higher than AA for c=3c=3. However, by applying Eq. 4, we get the new scores for proposal AA as [0.0,0.31,0.38,0.31][0.0,0.31,0.38,0.31] and for proposal BB as [0.55,0.0,0.27,0.18][0.55,0.0,0.27,0.18] — proposal AA is now ranked higher than BB for c=3c=3, leading to a higher AP for this class. As will be seen in Section 4, such a “re-ranking” property is the key to making NorCal excel in AP for all classes as well as in other metrics like APFixed{}^{\text{Fixed}} .

3 Comparison to Existing Work

Li et al. investigated classifier normalization for post-processing calibration. They modified the calculation of ϕc\phi_{c} from wc⊤fθ(x)\bm{w}_{c}^{\top}f_{\bm{\theta}}({\bm{x}}) to wc⊤∥wc∥2γfθ(x)\frac{\bm{w}_{c}^{\top}}{\|\bm{w}_{c}\|_{2}^{\gamma}}f_{\bm{\theta}}({\bm{x}}), building upon the observation that the classifier weights of head classes tend to exhibit larger norms . The results, however, were much worse than their proposed cost-sensitive method BaGS. They attributed the inferior result to the background class, and had combined two models, with or without classifier normalization, attempting to improve the accuracy. Our decomposition in Eq. 3 suggests a more straightforward way to handle the background class. Moreover, NcN_{c} provides a better signal for calibration than ∥wc∥2\|\bm{w}_{c}\|_{2}, according to . We provide more discussions and comparison results in the supplementary material.

4 Extension to Multiple Binary Sigmoid Classifiers

Many existing models for long-tailed object detection and instance segmentation are based on multiple binary classifiers instead of the softmax classifier . That is, scs_{c} in Eq. 2 becomes

in which wc\bm{w}_{c} treats every class c′≠cc^{\prime}\neq c and the background class together as the “negative” class. In other words, the background logit ϕC+1=wC+1⊤fθ(x)\phi_{C+1}=\bm{w}_{C+1}^{\top}f_{\bm{\theta}}({\bm{x}}) in Eq. 2 is not explicitly learned.

Our post-processing calibration approach can be extended to multiple binary classifiers as well. For example, Eq. 4 becomes

We note that solely calibrating the scores can re-rank the detected tuples across classes within each image such that rare and common objects, which initially have lower scores, could be included in the cap to largely increase the recall. Therefore, as will be shown in the experimental results, the improvement for multiple binary classifiers mainly comes from the rare and common objects.

However, one drawback of the score calibration alone is the infeasibility of normalization across classes; scs_{c} does not necessarily sum to 1, making it hard to re-order the scores of tuples within each class. Forcing the confidence scores across classes of each proposal to sum to 1 would inevitably turn many background patches into foreground proposals due to the lack of the background logit ϕC+1\phi_{C+1}.

Experiments

Dataset. We validate NorCal on the LVIS v1 dataset , a benchmark dataset for large-vocabulary instance segmentation which has 100K/19.8K/19.8K training/validation/test images. There are 1,203 categories, divided into three groups based on the number of training images per class: rare (1–10 images), common (11–100 images), and frequent (>>100 images). All results are reported on the validation set of LVIS v1. For comparisons to more existing works and different tasks, we also conduct detailed experiments and analyses on LVIS v0.5 , Objects365 , MSCOCO , and image classification datasets in the supplementary material.

Evaluation metrics. We adopt the standard mean Average Precision (AP) for evaluation. The cap over detected objects per image is set as 300300 (cf. Section 3.1). Following , we denote the mean AP for rare, common, and frequent categories by APr\text{AP}_{r}, APc\text{AP}_{c}, and APf\text{AP}_{f}, respectively. We also report results with a complementary metric APFixed\text{AP}^{\text{{Fixed}}} , which replaces the cap over detected objects per image by a cap over detected objects per class from the entire validation set. Namely, APFixed\text{AP}^{\text{{Fixed}}} removes the competition of confidence scores among classes within an image, making itself category-independent. We follow to set the per-class cap as 10,00010,000. Instead of Mask AP, we also report the results in APFixed\text{AP}^{\text{{Fixed}}} with Boundary IoU, following the standard evaluation metric in LVIS Challenge 2021https://www.lvisdataset.org/challenge_2021.. Meanwhile, we report APb{\text{AP}^{\textit{b}}}, which assesses the AP for the bounding boxes produced by the instance segmentation models.

Implementation details and variants. We apply NorCal to post-calibrate several representative baseline models, for which we use the released checkpoints from the corresponding papers. We focus on models that have a softmax classifier or multiple binary classifiers for assigning labels to proposalsSeveral existing methods (e.g., ) develop specific classification rules to which NorCal cannot be directly applied.. For NorCal, (a) we investigate different mechanisms by applying post-calibration to the classifier logits, exponentials, or probabilities (cf. Eq. 4); (b) we study different types of calibration factor aca_{c}, using the class-dependent temperature (CDT) presented in Eq. 5 or the effective number of samples (ENS) ; (c) we compare with or without score normalization. We tune the only hyper-parameter of NorCal (i.e., in aca_{c}) on training data.

2 Main Results

NorCal effectively improves baselines in diverse scenarios. We first apply NorCal to representative baselines for instance segmentation: (1) Mask R-CNN with feature pyramid networks , which is trained with repeated factor sampling (RFS), following the standard training procedure in ; (2) re-sampling/cost-sensitive based methods that have a multi-class classifier, e.g., cRT ; (3) re-sampling/cost-sensitive based methods that have multiple binary classifiers, e.g., EQL ; (4) data augmentation based methods, e.g., a state-of-the-art method MosaicOS . Please see the supplementary material for a comparison with other existing methods.

Table 1 provides our main results on LVIS v1. NorCal achieves consistent gains on top of all the models of different backbone architectures. For instance, for RFS with ResNet-50, the overall AP improves from 22.58%\% to 25.22%\%, including ∼7%/3%\sim 7\%/3\% gains on APr/APc\text{AP}_{r}/\text{AP}_{c} for rare/common objects. Importantly, we note that NorCal’s improvement is on almost all the evaluation metrics (columns), demonstrating a key strength of NorCal that is not commonly seen in literature: achieving overall gains without sacrificing the APf\text{AP}_{f} on frequent classes. We attribute this to the score normalization operation of NorCal: unlike which only re-ranks scores across categories, NorCal further re-ranks the scores within each category. Indeed, the only performance drop in Table 1 is on frequent classes for EQL, which is equipped with multiple binary classifiers such that score normalization across classes is infeasible (cf. Section 3.4). We provide more discussions in the ablation studies.

Comparison to existing post-calibration methods. We then compare our NorCal to other post-calibration techniques. Specifically, we compare to those in on the LVIS v1 instance segmentation task, including Histogram Binning , Bayesian binning into quantiles (BBQ) , Beta calibration , isotonic regression , and Platt scaling . We also compare to classifier normalization (τ\tau-normalized) on the LVIS v0.5 object detection task. All the hyper-parameters for calibration are tuned from the training data.

Table 3 shows the results. NorCal significantly outperforms other techniques on both tasks and can improve AP for all the classes. We attribute the improvement over methods studied in to two reasons: first, NorCal has only one hyper-parameter, while calibration methods in have hyper-parameters for every category and thus are sensitive to tune; second, NorCal performs score normalization, while does not. Compared to , the use of per-class data count in NorCal has been shown to outperform classifier norms for calibrating classifiers .

3 Ablation Studies and Analysis

We mainly conduct the ablation studies on the Mask R-CNN model (with ResNet-50 backbone and feature pyramid networks ), trained with repeated factor sampling (RFS) .

Effect of calibration mechanisms. In addition to reducing the logits, i.e., scaling down their exponentials (i.e., exp⁡(ϕc(x))/ac\exp(\phi_{c}({\bm{x}}))/a_{c} in Eq. 4), we investigate another two ways of score calibration. Specifically, we scale down the output logits from the network (i.e., ϕc(x)/ac\phi_{c}({\bm{x}})/a_{c}) or the probabilities from the classifier (i.e., p(c∣x)/acp(c|{\bm{x}})/a_{c}). Again, we keep the background class intact and apply score normalization. In Table 3, we see that scaling down the exponentials and probabilities perform the sameWith class score normalization, they are mathematically the same. and outperform scaling down logits. We note that, logits can be negative; thus, scaling them down might instead increases the scores. In contrast, exponentials and probabilities are non-negative, scaling them down thus are guaranteed to reduce the scores of frequent classes more than rare classes.

Effect of calibration factors aca_{c}. Beyond the class-dependent temperature (CDT) presented in Eq. 5, we study an alternative factor, inspired by the effective number of samples (ENS) . Specifically, we study ac=(1−γNc)/(1−γ)a_{c}=(1-\gamma^{N_{c}})/(1-\gamma) with γ∈[0,1)\gamma\in[0,1). Same as CDT, ENS has a single hyper-parameter γ\gamma that controls the degree of dependency between aca_{c} and NcN_{c}. If γ=0\gamma=0, we recover the original confidence scores. We report the comparison of these two calibration factors in Table 3. With appropriate post-calibration mechanisms, both provide consistent gains over the baseline model.

Importance of score normalization. Again in Table 3, we compare NorCal with or without score normalization across classes. That is, whether we include the denominator in Eq. 4 or not. By applying normalization, we see that NorCal can improve all categories, including frequent objects. Moreover, it is applicable to different types of calibration mechanisms as well as calibration factors. In contrast, the results without normalization degrade at frequent classes and sometimes even at common and rare classes. We attribute this to two reasons: first, score normalization enables the detected tuples of each class to be re-ranked (cf. Figure 3); second, with the background logits in the denominator, the calibrated and normalized scores can effectively prevent background patches from being classified into foreground objects. Please be referred to the supplementary material for additional results and ablation studies on sigmoid-based detectors (i.e., BALMS and RetinaNet ).

How to handle the background class? NorCal does not calibrate the background class logit. We ablate this design by multiplying the exponential of the background logit with a background calibration factor β\beta, i.e., exp⁡(ϕC+1(x))×β\exp(\phi_{C+1}({\bm{x}}))\times\beta. If β=1\beta=1, there is no calibration on background class. Figure 5 shows the average precision and recall of the model with NorCal w.r.t different β\beta. We see consistent performance for β≥1\beta\geq 1. For β<1\beta<1, the average precision drops along with reduced β\beta, especially for the rare classes whose average recall also drops. We note that, in the extreme case with β=0\beta=0, the background class will not contribute to the final calibrated score. Thus, many background patches may be classified as foregrounds and ranked higher than rare proposals. These results and explanation justifies one key ingredient of NorCal— keeping the background logit intact.

Sensitivity to the calibration factor. NorCal has one hyper-parameter: γ\gamma in the calibration factor aca_{c}, which controls the strength of calibration. We find that this can be tuned robustly on the training data, even on a 5K subset of training images: as shown in Figure 5, the AP trends on the training and validation sets at different γ\gamma are close to each other. In our experiments, we find that this observation applies to different models and backbone architectures.

NorCal reduces false positives and re-ranks predictions within each class. In Table 4, we show that NorCal can improve the AR for all classes but frequent objects (with a slight drop). The gains on AP for frequent classes thus suggest that NorCal can re-rank the detected tuples within each class, pushing many false positives to have scores lower than true positives.

NorCal is effective in APFixed\text{AP}^{\text{Fixed}} and Boundary IoU . Table 5 reports the results in APFixed\text{AP}^{\text{Fixed}} and APFixed\text{AP}^{\text{Fixed}} with Boundary IoU. We see that NorCal is metric-agnostic and can consistently improve the baseline model in all groups of categories. It suggests that the improvements are due to both across-class and within-class re-ranking.

Limiting detections per image. Finally, we evaluate NorCal by changing the cap on the number of detections per image. Specifically, we investigate reducing the default number of 300. The rationale is that an image seldom contains over 300 objects. Indeed, each LVIS image is annotated with around 12 object instances on average. We note that, to perform well in a smaller cap requires a model to rank most true positives in the front such that they can be included in the cap. In Figure 6, NorCal shows superior performance against the baseline model under all settings. It is worth noting that NorCal achieves better performance even using a strict 100 detections per image than the baseline model with 300.

Qualitative results. We show qualitative bounding box results on LVIS v1 in Figure 7. We compare the ground truths, the results of the baseline, and the results of NorCal. NorCal can not only detect more objects from the rare categories that may be overlooked by the baseline detector, but also improve the detection results on frequent objects. For instance, in the upper example of Figure 7, NorCal discovers a rare object “sugar bowl” without sacrificing any other frequent objects. Moreover, NorCal can improve the frequent classes, as shown in the bottom example of Figure 7. Please see the supplementary material for more qualitative results.

Conclusion

We present a post-processing calibration method called NorCal for addressing long-tailed object detection and instance segmentation. Our method is simple yet effective, requires no re-training of the already trained models, and can be compatible with many existing models to further boost the state of the art. We conduct extensive experiments to demonstrate the effectiveness of our method in diverse settings, as well as to validate our design choices and analyze our method’s mechanisms. We hope that our results and insights can encourage more future works on exploring the power of post-processing calibration in long-tailed object detection and instance segmentation.

Acknowledgments and Funding Transparency Statement

This research is partially supported by NSF IIS-2107077 and the OSU GI Development funds. We are thankful for the generous support of the computational resources by the Ohio Supercomputer Center. We thank Zhiyun Lu (Google) for feedback on an early draft of this paper and Han-Jia Ye (Nanjing University) for the help on image classification experiments.

References

Appendix A Additional Discussion on Related Work

Existing works can be categorized into re-sampling, cost-sensitive learning, and data augmentation.

Re-sampling changes the training data distribution — by sampling rare class data more often than frequent class ones — to mitigate the long-tailed distribution. Re-sampling is widely adopted as a simple but effective baseline approach . For example, repeat factor sampling (RFS) sets a repeat factor (i.e., sampling frequency) for each image based on the rarest object within that image; class-aware sampling samples a uniform amount of images per class for each mini-batch. Since an image can contain multiple object classes, Chang et al. proposed to re-sample on both the image and object instance levels. RFS is the baseline approach used for the LVIS dataset .

Cost-sensitive learning is the most popular category, which adjusts the cost of mis-classifying an instance or the loss of learning from an instance according to its true class label. Re-weighting is the simplest method of this kind, which gives each instance a class-specific weight in calculating the total loss (usually, tail classes with larger weights). The equalization loss (EQL) and EQL v2 ignore the negative gradients for rare class classifiers or equalize the positive-negative gradient ratio for each class to balance the training, respectively. The drop loss improves EQL by specifically handling the background class via re-weighting. The seesaw loss proposes a re-weighting scheme by combining the dataset statistics and training dynamics. Forest R-CNN leverages the class hierarchical for knowledge transfer and introduces new losses for hierarchical classification.

Instead of applying the new loss functions during the entire training phase, several recent methods decouple the training phase into two stages . At the first stage, the object detector is trained normally just like on a relatively balanced dataset such as MSCOCO . Then in the second stage, re-sampling or cost-sensitive learning is employed, usually for re-training or fine-tuning only the classification network. Such a pipeline is shown to learn both better features and classifier. For example, two-stage fine-tuning approach (TFA) first trains a base detector using only common and frequent classes, and then fine-tune the classifier and box regressor with re-sampling. Similar ideas are adopted in classifier re-training (cRT) , SimCal , balanced softmax (BSM) , balanced group softmax (BaGS) , DisAlign , and ACSL , which develop strategies or losses to re-train the classifier. Learning to segment the tail (LST) takes an incremental learning approach to gradually learn from the head to tail classes in multiple stages.

Data augmentation improves long-tailed object detection by augmenting data for the tail classes. DLWL and MosaicOS leveraged weakly-supervised data from YFCC-100M , ImageNet , and Internet to augment the long-tailed LVIS dataset . Copy-Paste self-augments the LVIS dataset by copying object instances from one image and paste to the others. Instead of augmenting images, FASA generates class-wise virtual features using a Gaussian prior whose parameters are estimated from features of real data.

A.2 Calibration of Model Uncertainty

We note that, the calibration rules we apply are different from the ones used for calibrating model uncertainty : we aim to adjust the prediction across classes, while the latter adjusts the predicted probability to reflect the true correctness likelihood. For calibrating model uncertainty, representative methods are Platt scaling , histogram binning , Bayesian binning into quantiles (BBQ) , isotonic regression , temperature scaling , beta and Dirichlet calibration , etc.

Appendix B Experimental Setups

Our approach NorCal is model-agnostic as long as the detector has a softmax classifier or multiple binary sigmoid classifiers for the objects and the background. Thus, we focus on those methods as long as the pre-trained models are applicable and public:

The baseline Mask R-CNN model with feature pyramid networks , which is trained with repeated factor sampling (RFS), following the standard training procedure in .

Re-sampling/cost-sensitive based methods that have a multi-class classifier for the foreground objects and the background class, e.g., cRT and TFA .

Re-sampling/cost-sensitive based methods that have multiple binary sigmoid-based classifiers, e.g., EQL , BALMS , and RetinaNet with focal loss .

Data augmentation based methods, e.g., MosaicOS . MosaicOS augments LVIS with images from ImageNet , which can improve the feature network of an object detector like Faster R-CNN or Mask R-CNN .

We note that, several methods change the decision/classification rules. For example, EQL v2 and Seesaw adopt a separate background or objectness branch during the training and inference. Some other methods (BaGS and Forest R-CNN ) re-organize the category groups and apply either a group-based softmax classifier or hierarchical classification. Therefore, it is not immediately obvious how to apply calibration to them.

B.2 Implementation

NorCal is easy to implement and requires no re-training of the model. We follow Eq. 4 and Eq. 5 of the main paper to apply NorCal to the existing models. For all the baseline detectors, we directly take the released models from the corresponding papers without any modifications. We report the results on the validation set with the best hyper-parameter tuned on training images for all models and benchmarks. The implementations are mainly based on the Detectron2 or MMdetection framework. We run our experiments on 4 NVIDIA RTX A6000 GPUs with AMD 3960X CPUs.

B.3 Inference and Evaluation

We follow the standard evaluation protocol for the LVIS benchmark . Specifically, during the inference, the threshold of confidence score is set to 10−410^{-4}, and we keep the top 300300 proposals as the predicted results. No test time augmentation is used. We adopt the standard mean Average Precision (AP) and denote the AP for rare, common, and frequent categories by APr\text{AP}_{r}, APc\text{AP}_{c}, and APf\text{AP}_{f}, respectively. For the object detection results on LVIS v0.5, we report the box AP for each category.

Appendix C Additional Experimental Results and Analyses

Due to space limitations, we only reported the results of NorCal with strong baseline models in the main paper (cf. Table 1). In this section, we provide detailed comparisons with more existing works on LVIS v1 and v0.5. We also examine NorCal on MSCOCO dataset . Moreover, we conduct further analyses and ablation studies of our method.

We summarize the results of instance segmentation on LVIS v1 in Table 6. As mentioned in Section B.1, several methods (e.g., BaGS , EQL v2 , Seesaw ) change the decision/classification rules and it is not immediately obvious how to apply calibration to them. Nevertheless, we include their results for comparison. We observe, for example, that NorCal can improve a simple baseline such as RFS to match or outperform all methods but Seesaw , which is trained with a stronger 2×\times schedule and an improved mask head. When paired with MosaicOS , NorCal can achieve state-of-the-art performance with all different backbone models, suggesting that improving the feature (especially on rare objects) and calibrating the classifier are key ingredients to the success of long-tailed object detection and instance segmentation.

C.2 Results on LVIS v0.5 Instance Segmentation

Many existing works focus on LVIS v0.5. In this subsection, we thus report the results of instance segmentation on LVIS v0.5 in Table 7. Again, we observe similar trends that NorCal can significantly improve the baseline models with all different backbone architectures. Particularly, we can also see improvements on the sigmoid-based object detector, i.e., BALMS .

C.3 Results on LVIS v0.5 Object Detection

In Table 8, we compare with existing methods that reported results on LVIS v0.5 object detection — only the bounding box annotations are used for model training. Concretely, we include EQL , LST , BaGS , TFA , and MosaicOS , as the compared methods. In addition, we study a popular sigmoid-based detector, i.e., RetinaNet with focal loss . We train the RetinaNet using the default hyper-parameters and apply NorCal on top of it. We see that NorCal can consistently improve the baseline models.

C.4 Results on Objects365 dataset

We further validate NorCal on Objects365 , a dataset designed to spur object detection research with a focus on diverse objects in the wild. Objects365 contains 2 million images, 30 million bounding boxes, and 365 categories with a long-tailed distribution. We train a Faster R-CNN as the baseline on the training set, with FPN and ResNet-50 as the backbone. We report results in Table 9. We not only show the overall mean AP, but also the mean APs for different groups of categories based on the training image number per category. NorCal outperforms the baseline detector on all groups of categories, justifying its effectiveness and generalizability.

C.5 Results on MSCOCO Dataset

We also experiment our method NorCal on the generic object detection benchmark, i.e., MSCOCO . MSCOCO is the most popular benchmark for object detection and instance segmentation, which contains 80 categories with a relative balanced class distribution (See Figure 8). More importantly, the least frequent class, “hair driver”, still has 189189 training images. In other words, all the classes in MSCOCO are considered as frequent classes using the definition of LVIS. We report results in Table 10. We see that the performance gains brought by NorCal is marginal. Our hypothesis is that the detectors trained with MSCOCO already see sufficient examples for all categories (even for tail classes) and the trained classifier is less biased.

C.6 Results on Image Classification Datasets

Besides object detection and instance segmentation, we further evaluate NorCal on three imbalanced classification benchmarks: ImageNet-LT , iNaturalist (2018 version) , and CIFAR-10-LT (with an imbalance factor 100) . ImageNet-LT has 1,000 classes while iNaturalist has 8,142 classes. All three datasets have long-tailed distributions on the number of training images per class but have a balanced evaluation set. We follow the literature to train a ResNet-50 classifier for the first two datasets, and a ResNet-32 classifier for CIFAR. Since there is no background class in these datasets, we simply drop the background class in Eq.4 in the main text. Results are shown in Table 11. As expected, NorCal consistently outperforms the baseline classifiers, demonstrating its effectiveness on long-tailed classification problems as well.

As mentioned in the Section 1 in the main paper, post-processing calibration for imbalanced or long-tailed classification has been studied in several prior works. Our approach is indeed inspired by their efficiency and effectiveness in classification problems and we extend them to the detection and instance segmentation problems.

C.7 Ablation Studies on Sigmoid-Based Detectors (i.e., with Multiple Binary Classifiers)

As shown in the main paper (cf. Table 3), we conduct ablation studies of NorCal with a standard softmax-based object detection. Here, we further examine a sigmoid-based object detector, i.e., BALMS , and report the results in Table 12. Beyond Eq. 7 of the main paper, we ablate NorCal with different calibration mechanisms, factors, and with and without score normalization. We note that, in this kind of models, CC binary classifiers are learned, each corresponds to one foreground class. In other words, no background class is specifically learned. Thus, the score normalization is usually not necessary or harmful — the background patches with low scores by all the classifiers will now gets their scores boosted due to calibration.

C.8 Empirical Class Frequency is Better than Classifier Norms for NorCal

As mentioned in the main paper (cf. Section 3.3 and Table 2 (bottom)), class-dependent temperature (NcγN_{c}^{\gamma}) provides a better signal for calibration than the classifier norms (∥wc∥2γ\|\bm{w}_{c}\|_{2}^{\gamma}) of the classifier. Table 13 shows a comparison between those two factors for our proposed calibration mechanism. With NorCal, we see that NcN_{c} outperforms ∥wc∥2\|\bm{w}_{c}\|_{2} on all object categories. Moreover, we notice that leaving the background intact shows a better performance, justifying our analysis and experimental results on how to handle the background class (cf. Section 3.2 and Figure 4 of the main paper).

C.9 Further Analysis on Existing Post-Processing Calibration Methods

We compare NorCal to the existing post-calibration methods in the main paper (cf. Table 2 (upper)). In the main paper, we follow the implementations in to perform the calibration after the top 300300 predicted boxes are selected. Here we study an alternative of directly applying the calibration before selecting the 300300 predictions. We show the results in Table 14. NorCal still outperforms all existing calibration methods.

C.10 Additional Qualitative Results

We provide additional qualitative results on LVIS v1 in Figure 9. We show the (predicted) bounding boxes from the ground truth annotations, the baseline Mask R-CNN with RFS , and NorCal.