Learning Open-World Object Proposals without Learning to Classify
Dahun Kim, Tsung-Yi Lin, Anelia Angelova, In So Kweon, Weicheng Kuo
Introduction
Object proposals are a set of regions or bounding boxes that contain objects with high likelihood . They have become the integral pre-processing steps for many computer vision systems, including object detection , segmentation , object discovery , weakly supervised object detection , visual tracking , content-aware retargeting , etc. Due to the success of object detection, the recent trend in object proposal research has shifted from object discovery to detection. While the goal of object discovery proposals is to propose any objects in the image, the goal of detection proposals is to propose only the labeled categories for downstream classifier. Learning-based proposals are popular detection proposals because of simplicity and shared computation with downstream detection. However, unlike their learning-free counterparts , these methods tend to overfit to annotated categories and struggle with novel objects . We want to ask, is it possible to combine the best of both worlds and “learn open-world (novel) object proposals?” This could potentially unlock learning-based proposals for promising applications including open-world detection / segmentation , robot grasping , egocentric video understanding , and large vocabulary detection .
Given a set of object annotations, we want to learn what general objects look like and propose highly dissimilar object candidates from unseen categories and new data sources. This matches the ability of human to detect novel objects in new environments without naming their categories e.g. a piece of obstacle on the road, a novel product on the shelf. Our main insight is that the classifiers in existing object proposers or class agnostic detectors impedes such generalization, because the model tends to overfit to labeled objects and treat the unlabeled objects in the training set as background. We propose Object Localization Network (OLN), which learns to detect objects by predicting how well a region is localized instead of performing foreground-background classification. This simple idea allows the model to learn stronger objectness cues. To our best knowledge, we are the first to demonstrate the value of learning pure localization-based objectness for proposing novel objects, although the idea of incorporating the localization quality estimation has been proposed by others in the standard fixed-category detection setting . We show that a classifier-free object proposer is key to achieve optimal cross-category and cross-dataset generalization, which is an important design difference to existing proposers or class agnostic detectors.
We study the efficacy of OLN on the COCO cross category setting following existing works . Despite the simplicity, OLN outperforms the state-of-the-art by +3.3 AUC (+5.0 AR@10, +5.1 AR@100) on novel categories. Our ablation studies confirm that the use of foreground-vs-background classifier hurts, and that localization helps. In addition, we study cross-dataset generalization from COCO to RoboNet , Objects365 , and EpicKitchens . We chose RoboNet because it contains a wide range of novel objects common in robotics grasping application and the bin environment permits more reliable exhaustive annotation for proper evaluation . On RoboNet, OLN performs exhaustive, class-agnostic object detection and outperforms the standard approach by +1316 AP, whereas on Objects365 OLN has +4 AR@10 and +8 AR@100 over the standard approach. Qualitative visualization on EpicKitchens further shows that OLN outperforms the standard approach in detecting a variety of novel objects. Last but not least, we apply OLN as a drop-in replacement for RPN on LVIS long tail detection and observe a gain of +1.4 AP, mostly attributed to the rare (+3.4 APr) and common categories (+1.8 APc). This shows OLN is able to capture the long tail in large vocabulary detection.
It is worth noting that estimating localization quality is not new in the standard detection, but they are always used alongside classification and validated on seen categories only, e.g. FCOS . To our knowledge, we are the first to explore the use of localization cues independent of classification for object proposals. This discovery helps us obtain notable gains on COCO and generalize to many dissimilar datasets better than existing method.
Our contributions can be summarized as follows:
To our knowledge, we are the first to show the value of pure localization-based objectness learning for novel object proposals, and propose a simple-yet-effective classifier-free Object Localization Network (OLN).
Our approach outperforms state-of-the-art methods on cross-category setting on COCO and improves cross-dataset settings on RoboNet and Object365, long-tail detection (LVIS) and egocentric videos (EpicKitchens) over the standard approach.
We carefully annotated the RoboNet dataset for the presence of all objects in an exhaustive fashion. We perform open-world class-agnostic object detection, and evaluate the Average Precision, which also improves existing AR-based evaluation of proposals on partially-annotated data.
Extensive ablation and analysis on OLN modeling choices reveal the benefits of each localization cue and the overfitting of existing classifier-based methods.
Related Work
Below we discuss existing efforts to improve proposal and detection quality, and works that scale detection to more visual categories.
Object proposal. In early works, the emphasis was on category-independent object proposals , where the goal is to identify instances of all objects in the image irrespective of their category. These works utilize hand-crafted heuristics to capture the notion of general objects , i.e., color contrast, edge.
Recently, learning-based proposals have demonstrated better performance than classical approaches in both precision and recall, and are an important part of two-stage detectors. A representative example is region proposal network (RPN) which identifies a set of regions in a given image that could contain objects, which are then used by a downstream detector module to localize and classify objects. A number of follow-up works have been proposed to improve the quality of such region proposals and reduce their number in order to speed up the final detection task. In fact, these proposal modules are trained end-to-end with the detector module, where the notion of objectness is defined by a set of training categories in the dataset. Despite their progress in detecting objects of the known, supervised categories, learning-based proposals still struggle on novel objects.
More closely related to our work are studies on generalization of object proposals onto unseen classes. Chavali et al. demonstrate the standard proposal evaluation is problematic or “gameable” in evaluating category-independence of object proposal because detection of unknown object classes is explicitly penalized in the benchmark protocol. Wang et al. study the generalization from a dataset standpoint and demonstrate the impact of visual diversity and label granularity of training dataset on the generalization of object proposers. In contrast, we focus on the modeling choices in designing an object proposer that can generalize to novel categories and new datasets.
Multi-class detection. A number of efforts have been made to scale up the number of classes for detection by transferring commonalities between object categories with varying degrees and quality of supervision.
Weakly-supervised approaches aim to utilize abundant image-level labels and leverage class-agnostic box proposals to build detectors. Approaches under semi-supervised setup employ the weak image-level labels for novel classes as well as box-level labels for base classes. For example, YOLO-9000 and R-FCN-3000 concurrently train on box-level and image-level data to scale up the detector’s class coverage. Knowledge transfer-based methods learn to transfer the proposals from base to novel classes based on their similarity on semantic hierarchy. This line of research is also related to few-shot and zero-shot detection methods, which attempt to detect novel classes given only a few samples or class descriptions.
In contrast to the above class-specific detectors, our goal is to go beyond the concept of category and detect all objects in a category-independent manner (without classification). Even though the multi-class detectors enumerate many categories, they still fail to generalize for unseen/unknown object categories.
Existing datasets focus on single-dataset settings . Models trained on one dataset are only evaluated on the same dataset. Recently the Robust Vision Challenge is a first step towards cross-dataset benchmark. We hypothesize that learning objectness on one dataset should also transfer to another dataset, as the saliency cues tend to be more generalizable than class-specific information. Therefore, in this work we study the generalization setting of training on COCO, then testing on other datasets: RoboNet , Objects365 , EpicKitchens , and LVIS .
Researchers have studied many ways to improve the localization quality in object detections by learning centerness , iterative proposal refinement , or box/mask IoU prediction . These approaches have demonstrated meaningful improvement in the standard object detection task. However, it remains an open question whether these methods can transfer to novel categories. In addition, most of these works use these localization cues alongside classifier outputs. It is unclear whether these cues on their own can effectively distinguish objects from background.
Proposed Method
Before we describe OLN in section 3.2 and 3.3, we would like to define the baselines which can address the same unseen category generalization problem. Region proposal network (RPN) is the most common family of approaches of objectness learning in object detection. By design, RPNs aim to propose all objects in the image regardless of their categories, but in practice they often struggle when encountered with novel objects in the open world. Another family of baseline is to train existing object detectors in a class-agnostic fashion by treating all annotated categories as one foreground category. As OLN is built upon RPN and Faster R-CNN, we use both as strong baselines throughout the paper. Moreover, we provide comparisons with different state-of-the-art models in region proposal and object detection . As seen in our experiments later, OLN outperforms all of them in various generalization scenarios.
2 Pure localization-based objectness
In the context of learning-based object proposals, “object” is defined as a set of annotated categories, and the learning of objectness is cast as a binary classification task: whether or not a region belongs to the union of the pre-defined categories. However, our main insight is that such discriminative learning of the foreground-vs-background problem impedes generalization because the model learns to classify the unlabeled/unknown objects as background. To address this problem, we propose a non-discriminative and classification-free notion of objectness.
The classification view of “objectness” is to ask “how much does this region look like a foreground object?” From a localization standpoint, we want to ask instead “how well does this region overlap with any ground-truth object?”. Our intuition is that every object can be characterized by its location and shape, regardless of its category. OLN leverages these geometric cues to capture the objectness of a proposed region. We demonstrate that the learnt objectness cues based on the localization (location and shape) quality can collectively improve the generalization of object proposals beyond the labeled categories and data sources. We adopt centerness and IoU score for location and shape quality measures respectively, while not restricting other choices such as Dice coefficient and generalized-IoU .
The idea of incorporating localization quality is not totally new in object detection. Several works recalibrate the final detection confidence by using both localization and classification subnets. Note that, however, their localization cues are thoroughly an auxiliary to the classifiers and are devised for within-category detection. In contrast, we demonstrate that a pure localization-based objectness is the key to generalize beyond category and across datasets, and that a classification head severely hurts generalization. To best our knowledge, this intuition has not been discussed in any of the prior works.
3 Object Localization Network (OLN)
The goal of OLN is to learn localization for objects and enable better generalization to new and unseen categories. OLN is a two-stage object proposer (see Figure 2). Similar to Faster R-CNN , OLN consists of a fully-convolutional FCN stage and region-based RoI stage, but the key difference is that the classifiers in both FPN and ROI stages are replaced with localization quality predictions.
The input to this region proposal stage is the features from each level of ResNet feature pyramid . Each feature map goes through a convolution layer followed by two separate layers, one for bounding box regression and the other for localization quality prediction. The network architecture design follows from the standard RPN heads.
We choose centerness as the localization quality target and train both heads with L1 losses. Learning localization instead of classification at the proposal stage is crucial as it avoids overfitting to the foreground by classification. For training the localization quality estimation branch, we randomly sample 256 anchors having an IoU larger than 0.3 with the matched ground-truth boxes, without any explicit background sampling. For the box regression, we replace the standard box-delta targets (xyhw) with distances from the location to four sides of the ground-truth box (lrtb) as in . We choose to use one anchor per feature location as opposed to 3 in RPN, because we observe its better generalization as each anchor can ingest more data.
We take the top-scoring (e.g., well-centered) proposals from OLN-RPN and perform RoIAlign to extract the region features from each feature pyramid level. Then we linearize each region features and feed it through two fc layers, followed by two separate fc layers, one for bounding box regression and the other for localization quality prediction. We use the same network architecture as Faster R-CNN heads . We choose IoU as the localization quality target and train both heads with L1 losses. Learning localization quality at the second stage is integral as it allows the model to refine the proposal scoring and simultaneously avoid overfitting to the foreground. Compared to IoU-Net which requires manual proposal generation for IoU training, OLN directly computes the IoU targets from the OLN-RPN proposals and ground-truth boxes, thus greatly saving computation costs.
Extension - OLN-Mask. We explore whether more localization learning can further improve the generalization ability of our framework. To this end, we extend our OLN-Box model to perform mask prediction by adding the class-agnostic FCN mask head of Mask R-CNN, which we refer to as OLN-Mask model. Following our philosophy of OLN, and similarly to MS R-CNN , we learn to regress the IoU between the predicted and its GT mask.
Our mask-IoU predictor directly branches out from the fourth layer of the added FCN mask head, without having a feedback connection from the mask prediction unlike . The IoU branch consists of a convolution layer, a max pooling layer and three fully connected layers. During training, we assume mask annotations are available for the training categories, and use smooth-L1 loss for IoU regression.
During inference, the objectness score of a region is computed as a geometric mean of the centerness and IoU scores (box:, mask:) estimated by OLN-RPN and OLN-Box branches. For OLN-Box, the score . For OLN-Mask, the score is .
Experiments
We study the generalization ability of the learned object proposal networks in three challenging generalization scenarios. We outperform all methods and strong baselines on them, and in some cases, significantly. First, we study 1) cross-category generalization on COCO dataset by evaluating average recall (AR) on new unseen classes. We compare with other state-of-the-art proposal and detection methods, and provide extensive ablation and analysis on OLN design. Also, we explore more challenging 2) open-world class-agnostic detection where the setup requires to detect all objects in an exhaustive and class-agnostic fashion. The testing images contain highly dissimilar objects to those in the training dataset; We train a model on COCO and test on RoboNet dataset and evaluate the Average Precision (AP). We demonstrate more 3) cross-dataset generalization results by testing on Object365 and EpicKitchens datasets. Finally, we study the 4) impact on long-tail object detection of OLN on LVIS dataset. In the following, we provide experiment setups, evaluation protocol and results for each setting. In each §4.1, §4.2, §4.3, and §4.4, we use the same training and testing pipelines for all competing methods, for a fair comparison. All experiments use the same ResNet-50 with feature pyramid backbone, and the same box regression head and training hyper-parameters as the Faster R-CNN , unless specified otherwise. More details are in supplementary materials.
We split the COCO dataset into 20 seen (VOC) classes and 60 unseen (non-VOC) classes. We train a model with box annotations of only seen classes, and evaluate the recall on unseen non-VOC classes only. To avoid evaluating any recall on seen-class objects, we do not count those seen-class detection boxes into the budget when computing the Average Recall (AR@k) scores.
We compare with learning-based single-stage and multi-stage methods. The comparison is in Table 1. For all methods, we use the standard official models available in MMDetection and train with the same default training and testing pipeline. Batch size of 16 (2 per GPU) and an initial learning rate of 0.02 are used. Note that the detection models are their class-agnostic versions with a binary classifier. The NMS threshold is set to 0.7.
OLN-RPN improves the standard RPN by a large margin of +5.2 AUC (+4.3 AR10 and +7.4 AR100), and outperforms all single-stage competitors. We also include RPN trained with 1-or-0 linear regression L1 loss instead of cross entropy loss and show that simply changing the loss function does not help. FCOS with a binary classifier is a popular example of combining classification and localization cue (centerness). Two-stage OLN-Box shows a large gain of +4.9 AUC (+6.1 AR10 and +7.3 AR100) over Faster R-CNN. Guided anchoring (GA-RPN) is an advanced iterative RPN with deformable convolution, and Cascade RPN is the state-of-the-art RPN method with multi-stage refinement. OLN-Box outperforms all of them with a healthy margin: +3.3 AUC. In terms of model size, OLN-RPN is the same size as RPN, and OLN-Box is the same as Faster R-CNN.
We can come up with a stronger Faster R-CNN baseline where we do not suppress unseen classes in the classification. We filter out all unseen class objects to train Faster R-CNN *filtered* (BG sampling at IoU = 0 with unseen classes). Table 1 shows that OLN does better by learning to localize (+3.0 AUC), even when Faster R-CNN *filtered* has access to the ground truth boxes of novel objects.
It’s worth noting that this paper’s focus is on objectness learning for generalization of proposal models; we do not explore orthogonal factors such as advanced convolutions for feature alignment and the use of box regression statistics , that may further improve performance.
We borrowed the DeepMask setting that evaluates on ‘all’ categories in Table 2. To make a cleaner comparison, we report in Table 1 the MCG (SOTA learning-free baseline) on ‘unseen’ split and find that OLN still maintains a healthy margin above it.
We study what modeling component helps or hurts the generalization of OLN. We enumerate over different objectness cues, i.e., classification, centerness, IoU and Dice score, and their single-stage and two-stage configurations. The ablation is in Table 3 and Table 4.
We notice that the second refinement head of two-stage models lead to better performance throughout all choices of objectness measure (Table 3 - vs -, vs -, and vs -). This can be attributed to better bounding box regression which has additional layers following the first stage, as also noted by He et al. . This trend can be also seen between other single-stage vs multi-stage methods in Table 1.
We perform ablations on what objectness cue hurts or helps each stage of OLN. We start with single-stage models (-) where the box regression pipeline is analogous to that of standard RPN except the transformed box coordinates. We compare classification, centerness and IoU scores in Table 3; both localization-based scores outperform classification, and the centerness shows the best AR. Among the two-stage configurations, model-() is comparable to Faster R-CNN, where only the last class score is used at testing. Again, we observe the overall superiority of localization-based objectness learning. We also notice that the second stage prefers IoU learning over centerness learning, whichever one of them is used in the first stage ( vs and vs ). Intuitively, IoU measure is sensitive to both location and shape of the detection, and thus better captures the quality of variable-sized boxes in the second stage. On the other hand, centerness can be a more suitable measure for the first region proposal stage, where the shape of the anchors are fixed. Overall, the combination of centerness-then-IoU () shows the best performance, implying their own contributions to the generalization. We also demonstrate that a different choice of localization quality also performs well (). For the rest of the paper, OLN-RPN and OLN-Box refer to model-() and () respectively.
Table 4 shows the results by adding a classifier branch to the best-performing OLN-RPN () and OLN-Box () models. Note that all the used objectness scores are geometrically averaged at test time. Throughout the single-stage and two-stage configurations, we observe a consistent drop in AR when adding a binary classifier. Largest drops are seen when having the classifiers in both stages. This validates our hypothesis that discriminative learning of object-or-not classification impedes generalization of object proposals, and that a pure localization-based objectness is the key to generalization.
Is the gain of OLN from the fact that we sample around the objects and avoid penalizing the unlabeled background objects? or does the permissive definition for positive helps generalization? In Table 5 we run Faster R-CNN with low sampling of background anchors similar to OLN. We validate that changing the sampling ratio does not help the baseline. This is because a balanced ratio (e.g., 1:1) of explicit negative vs positive samples is required for the binary classifier. Table 5 also shows various definition of positive / negative samples and that OLN-Box still outperforms Faster R-CNN and *filtered* models by a healthy margin. We think it is becuase OLN can learn more from the positives using localization cues instead of binary classification.
In Table 6 we show the results for the mask-extended OLN model, OLN-Mask. We validate that incorporation of additional localization quality learning, i.e., mask-IoU, can further improve both the box-AR and mask-AR performances.
We visualize class-agnostic Mask R-CNN and our OLN-Mask trained on VOC categories on COCO in Figure 3. The images come from COCO validation set. We observe Mask R-CNN tends to miss out-of-VOC class objects even when they appear in the dominant and salient region in an image, e.g., refrigerator (left column) and giraffes (right column). We believe this is because classification leads to learning just the union of 20 VOC categories and use that knowledge to detect objectness. On the other hand, OLN can generalize well to detect many novel objects, such as fridge, microwave, vases, and giraffes, showing that localization leads to learning a more general notion of objectness.
In Figure 4, we visualize the confidence heatmap of different localization cues vs classifier. We can see the IoU and centerness cues generalize better to novel objects than classification. For example, both IoU and centerness capture the flying kite, the traffic lights, handicrafts, the child and his/her toys, while the classifier head misses them. We use model-, , and in Table 3 for classification, IoU and centerness visualization respectively.
2 Open-World Class-Agnostic Detection
Our ultimate goal is to learn a generalizable object detector that can detect every object in our open world. Taking a step towards this goal, we investigate the open-world, class-agnostic detection ability on highly diverse and hard-to-name objects in RoboNet dataset. We train a model on COCO and test on held-out RoboNet dataset.
Chavali et al. report the vulnerability of current evaluation protocol for object proposals when evaluating on a partially-annotated dataset such as COCO, because it fails to capture the performance (e.g., reward the recall) on all the non-COCO categories that are not annotated in the ground truth. To handle this issue, we exhaustively annotated bounding boxes for the presence of all objects in RoboNet images. This allows us to measure Average Precision as in standard object detection evaluation. Such evaluation captures whether a model can generate a few but highly-precise detections. We use NMS threshold of 0.5 and evaluate box-AP on top 100 proposals in this experiment.
To enable the evaluation of our task also by the standard AP metric, we carefully annotated the RoboNet dataset for the presence of all objects in an exhaustive fashion. RoboNet is a large-scale robot manipulation video dataset collected for pre-training reinforcement learning . The dataset contains very challenging clutter and object diversity that we find suitable for evaluating general objectness learning. Since the video data contains lots of redundancy at frame level, we resort to random sampling to construct a concise yet diverse evaluation set. Concretely, we randomly sampled 109 videos from the 162417 videos in the training set (roughly 1/1500) and for each video we select only the center frame. The parameters are picked by visually inspecting the diversity of resulting data (See Figure 6-(a, c) for examples). Our annotated dataset consists of total 109 frames, 1277 objects, and 34.8% of small objects, which are those with long side 32 pixels (same as COCO). For object bounding box annotation, we adopt the method of labeling extreme points as proposed in , which results in better accuracy and efficiency for the raters.
To exclude the background scene in the image which could contain hard-to-define objects/building structures, we label the object bin as well and crop out only the within-bin region for evaluation. This removes most of the ambiguity in objectness evaluation and allows us to have reasonably exhaustive annotation of objects for detection evaluation via average precision. To validate whether the size is large enough, we trained 5 independent runs of our model and observe a variance of 0.17 AP, showing that the results are quite reproducible on our dataset size.
In this rigorous evaluation setting, larger gains are revealed for OLN models than in other experiments. In Table 7 we compare COCO →RoboNet generalization performance of OLN-Box and OLN-Mask models and class-agnostic Faster R-CNN and Mask R-CNN baselines. OLN models greatly outperform RPN and Faster R-CNN by +12.7 and 15.7 points AP, respectively. We also observe that the improvement in APIoU=.50 is significantly larger, by +22.0 and 29.1 points AP. These results provide strong evidence that OLN can achieve both high precision and recall in detecting novel objects, and that a few generated OLN proposals (100) can be directly used as final bounding boxes for all objects in an image, i.e., class-agnostic detection.
Figure 5 visualizes the box and mask predictions from baseline class-agnostic Mask R-CNN vs our OLN-Mask model. OLN is able to detect many novel objects (e.g. gloves, toys, tape, parts of toys), while the baseline misses most of them. Figure 6 presents an analysis of OLN failure modes by comparison with the ground-truth boxes. The false positive example is where OLN detects part of an object, e.g., it detects handle and head of a cleaning brush as two individual objects.
3 More cross-dataset generalization
We further study the generalization ability of OLN from one data source to different data sources.
Objects365 is a large-scale object detection dataset consisting of 365 object categories from our daily lives, which are a super-set of 80 COCO categories. We train a model on COCO and test on non-overlapping classes on this dataset. We follow the same AR@k evaluation protocol in Section 4.1, treating all the 365 classes as a single ‘object’ class, and leaving out the detection boxes on seen COCO 80 classes when evaluating top-k recall.
Table 8 shows the COCO →Objects365 generalization results. OLN outperforms all the classification-based baselines for all number of proposals considered. Especially, the gains over the RPN and Faster R-CNN start from +4.5 and 3.3 AR10 and achieve large gaps of +8.2 and 7.6 points in AR100.
3.2 Generalization to EpicKitchens
EpicKitchens is the largest egocentric vision dataset containing the wearers’ different kitchen activities like washing and cooking, with many different objects like food, kitchenware and appliances. It provides the automatic mask annotations of 66M objects, which they extracted by using the off-the-shelf Mask R-CNN trained on COCO dataset. Therefore, we use their official mask annotations as our baseline in this experiment.
AR or AP are difficult to compute on EpicKitchens because there are no groundtruth boxes or masks. Therefore, we visually compare the EpicKitchens mask annotations and our OLN-Mask prediction outputs in Figure 1 and 7, which are randomly selected examples. We observe that OLN-Mask is able to detect almost complete set of objects in an image, while the baseline Mask R-CNN misses most out-of-sample objects. To illustrate, OLN is able to detect a stack of dishes, dough balls, a toaster, half side of a microwave and its front window, as shown in Figure 1 and 7. These are novel categories outside of the COCO vocabulary, and appear in different configurations and viewpoints rarely seen in COCO. This demonstrate the clear benefits of learning localization for novel object detection.
4 Impact on long-tail object detection
Large vocabulary detection has gained much attention lately in research community because the problem captures the Zipf’s law of visual world, i.e. most categories appear with low frequency. We want to see whether OLN can aid downstream detection in a challenging long-tail setting, on LVIS dataset , e.g., only 1-10 training samples are available for rare categories. Table 9 shows that a drop-in replacement of RPN with OLN-RPN helps the overall average recall (AR) by +1.5, where most gains come from rare (+5.3) and common categories (+1.9). The 5-point gain in rare categories recall demonstrates that OLN can benefit the long-tail of large vocabulary detection. In fact, this gain transfers quite well to overall average precision (AP) improvement +1.4, where most gains come from the rare (+3.4) and common (+1.8) categories. We choose detections per image, following the convention of the community. Our baseline matches the performance of Faster R-CNN reported by in the same setting.
Figure 8 presents a few challenging cases of long-tail detection on rare categories - piles of papaya, crayon, and a small roller-blade. Compared to the baseline, a drop-in OLN-RPN replacement enables better detection on these challenging cases. We can see many missed papayas, crayons and the small roller-blade are now correctly localized and classified.
Conclusion
In this paper we tackle the challenging problem of learning novel object proposals. Observing the tendency of existing proposals to overfit to training categories, we propose a simple yet effective framework (OLN) that learns to propose novel objects by learning localization cues (centerness, IoU, and regression) instead of binary classification. Experiments show that OLN outperforms existing methods on cross-category generalization on COCO, as well as cross-dataset settings on RoboNet, Object365, and EpicKitchens. Moreover, a drop-in replacement of OLN improves the performance on large vocabulary and egocentric video object detection on LVIS dataset. We believe OLN is a step forward in open-world novel object understanding with many applications.