Zero-Shot Detection
Pengkai Zhu, Hanxiao Wang, Venkatesh Saligrama
I Introduction
Zero shot learning (ZSL) has recently drawn increasing attention. By leveraging the inexpensive descriptions of categories, connections between the visual and semantic representations are built to interpolate the unseen class, which has no examples in the training stage. While methods like achieve high classification accuracy over unseen categories, a more realistic problem, generalized zero shot learning (gZSL), where both seen and unseen classes need to be recognized, is proposed. This problem makes ZSL more practical since there is no oracle in the real world indicating whether an object class has been seen during training.
However, this extension of the classical ZSL problem still has its own limitation: it assumes the object is precisely localized and the only task is to recognize it. In fact, there are many potentially unseen objects appearing in the wild. An intelligent system should be able to not only classify, but also to localize them. Therefore, in this paper, we consider this additional source of complexity, and introduce zero shot detection (ZSD). This involves detecting unseen objects.
Deep learning based object detection methods trained on fully annotated training data has had significant success over the last few years . Methods such as YOLOv2 , achieve high-performance, by training detectors to accurately predict object bounding boxes to match the ground-truth bounding boxes, while suppressing bounding boxes that are part of the background image.
Nevertheless, YOLOv2 as currently structured, is also somewhat limited in scope in the context of large-scale applications where we encounter a large number of object classes. In these applications, it is unrealistic to expect object bounding boxes for all classes at sufficient scale that is required for training. Indeed, the lack of labeled training data in large-scale recognition applications has led to the emergence of zero-shot recognition (ZSR) methods (see for instance, ) as an alternative means to supplement labeled data. As object detection moves towards large-scale Although the number of annotations have increased in common detection datasets (e.g. 20 classes provided by PASCAL VOC and 80 provided by the more recent MSCOCO ), it is substantially smaller relative to image classification ., it is imperative that we move towards a framework that serves the dual role of detecting objects seen during training as well as detecting heretofore unseen categories. Furthermore, this problem gains more importance as we move towards object detection appearing in the wild.
Motivated by these challenges, we develop a novel zero-shot detection architecture (ZS-YOLO) for detection of unseen object classes. Our method is based on a seamless integration of semantic attribute predictors with YOLOv2’s visual detection architectureWe choose YOLOv2 for concreteness. Our method is applicable to other architectures. See Sec.III.B. Specifically, we train an end-to-end model for zero-shot detection based on a novel multi-task loss objective, which incorporates semantic and visual information. Nevertheless, at test-time, our method is agnostic to semantic information of unseen objects, and the semantic component of our network functions as a system for identifying semantic components that resemble trained classes. We choose YOLOv2 as the base detector for zero-shot detection because it is the state-of-the-art single stage detector on existing benchmark datasets. By changing the confidence loss and network backbone, our method can be easily applied to other single stage detector like SSD and RetinaNet. In addition, ZS-YOLO can be viewed as a variation of region proposal network (RPN), thus can be integrated with two-stage detectors like Faster-RCNN seamlessly. Ultimately, our choice of YOLOv2 is coincidental based partially on the ease with which we can integrate other side information, and the fundamental focus of the paper is on understanding and quantifying the utility of semantic attributes for zero-shot detection.
Limitations of Naive Methods: In comparison various naive strategies do not perform well. For instance, a cascade of YOLOv2 and a off-the-shelf zero-shot classifier (ZSR) at run-time turns out to be somewhat less effective in detecting unseen classes. This may seem surprising particularly because ZSR is capable of transferring semantic class-level information (such as attributes , word phrases ) for synthesizing unseen object classifiers.
The fundamental issue is that, YOLOv2 tends to relegate unseen objects as background leading to missed detection of unseen objects. In our experiments we often observe significant recall drops when a detector is trained on seen classes but applied to unseen object classes with zero training samples. YOLOv2’s low recall rate for unseen objects is altogether not surprising because YOLOv2 achieves high-performance, in part, by implicitly learning unique object visual features that can be used to match true objects at test-time through penalization of objectiveness confidence (to suppress false positives and maximize precision) during training. As a result, if an unseen class does not share the visual features learned on training data, these unseen objects have low objectiveness scores resulting in low recall.
Key Insight: We posit that we must explicitly train detectors that account for semantic attribute information in the visual domain to ensure that semantically similar object attributes will be reflected in the learnt visual features. In this way, at test-time, we can hope to detect unseen objects that are semantically related to seen object classes.
Recognition vs. Object Detection: We focus on the problem of proposing object bounding boxes in presence of seen and unseen classes, although, our method can be readily extended to recognition if semantic information for unseen objects were available. We focus on object detection out of necessity (limitations of available datasets) as well as practical reasons (for objects in the wild scenarios semantic information for unseen classes is not available).
To train ZS-YOLO, we learn an end-to-end detection network with a hierarchical architecture (Fig. 2): In the first level, we train the network with a multi-task loss to perform (1) bounding box prediction in the visual domain; and (2) attribute prediction in the semantic domain; Next, the visual features of each bounding box proposal and its semantic attributes are further combined as a multi-modal input for the final layer to produce an objectiveness confidence score. This setup must be contrasted with existing detection frameworks which predict the confidence score based solely on the visual space. Extensive experiments are conducted to verify the ZS-YOLO performance on both traditional zero shot detection where only unseen objects exist and generalized zero shot detection (GZSD) where both seen and unseen objects appear in the images. Our experiments evidently shows the benefits of our multi-task training and multi-modal confidence prediction strategy: we improve the recall rate of the baseline model from to on unseen classes at 0.9 confidence, and unseen average precision from to on PASCAL VOC dataset for 10(seen)/10(unseen) split. Similar trends are also observed on different data splits as well as MS COCO dataset. We then perform extensive ablative analysis and quantify the utility of semantic information in for unseen object detection. We then identify cases where attribute information is particularly useful for unseen object detection.
Our contributions in this paper are: (1) Novel method for zero-shot detection problem that seamlessly integrates semantic attribute predictors with visual features during training. (2) Dataset: We construct a new ZSD dataset with multiple seen and unseen classes splits based on existing PASCAL VOC and MS COCO dataset; New performance metrics are also introduced and discussed; (3) We develop a new ZSD detector, based on the visual network structure of YOLOv2 . In contrast to state-of-art detectors, ZS-YOLO learns to predict semantic attribute as a side task during training, and produces object bounding boxes with both visual and semantic information. We observe significant improvements on both PASCAL VOC and MS COCO for unseen classes.
II Related Work
We utilize YOLOv2 as our baseline model, since it is fast ( FPS detection speed), simple (single-shot detection with a fully convolutional network structure ), and effective (state-of-art performance). As a single-shot detector, YOLOv2 directly infers from image cells and simultaneously produces a fixed set of bounding box proposals together with their associated confidence scores. Compared to YOLOv2, while other contemporary deep detectors can be as effective, they are less efficient (usually FPS). For instance, Faster-RCNN and R-FCN use a Region Proposal Network (RPN) as a parallel branch to first generate object proposals, and use the pooled region features to further refine bounding box locations as well as their objectness scores, which is usually much slower compared to YOLOv2. Nevertheless, in the context of large-scale detection problems, these methods including YOLOv2 are somewhat ineffective in detecting unseen object classes when trained with no corresponding training data. This can be attributed to seeking high supervised performance during training, a strategy that leads to maximizing detection precision with seen classes and encourages the network to suppress any image regions with divergent visual features as false positive proposals. This strategy in turn hampers unseen object detection since proposals of novel or unseen objects are often relegated as background.
On the other hand, in contrast to deep detectors, which are trained to suppress false positives, region proposal methods can potentially discover objects that are unseen. However, region proposal methods suffer from significantly high false positive rates. While they are designed to propose hundreds of regions per object, this leads to poor precision and requires significant computational resources for post-processing to improve accuracy. For these reasons we leverage YOLOv2 rather than improve precision of object proposal methods. Specifically, we extend YOLOv2 by leveraging the semantic attributes in the detection architecture. Nevertheless our method can be applied to other detectors, but we choose to demonstrate it on YOLOv2 due to superior performance.
II-B Zero Shot Recognition
Zero Shot Learning (ZSL) seeks to recognize novel visual categories that are unannotated in training data . Traditional ZSL is constrained to closed set of unseen classes only. Recently it has been shown that traditional methods do not generalize well to the case where classification includes both seen and unseen classes at test time . While generalized ZSL (GZSL) relaxes this constraint that the test data only belongs to unseen class, it still does not deal with background class which significantly increases the problem complexity. Our method complements the existing ZSL work by reducing the open set problem to a closed set one by detecting only the foreground objects. After bounding boxes of unseen objects are extracted by our detector, any ZSL method can be applied to obtain object labels.
II-C Other Methods
Recent work on weakly supervised localization have proposed methods to localize objects without bounding box annotations . However, these methods still rely on image level object annotations and are not focused on detecting the unseen. In our case, there are no annotations for unseen classes in training data. There are also methods that aim to discover objects without supervision . These methods usually rely on redundant parts and features in the training data to discover patterns and clusters of objects. Our goal is to transfer the detection knowledge of seen classes through semantic attributes to unseen ones. Fully unsupervised methods do not utilize the semantic transfer and may have hard time discovering clusters that are completely missing in the training data.
II-D Concurrent Works
Two other groups have concurrently worked on ZSD and appeared around the time a preliminary version of this paper was posted on arxiv. However, both approaches differ significantly from ours. Both works assume object proposals are predicted by a predefined proposal generator (Edge-Box for and Region Proposal Network(RPN) for ), and focus on the subsequent ZSR problem, which is to map the visual features extracted from object proposals to a semantic embedding and perform classification. Nevertheless, this assumption of already available proposals on unseen objects is unrealistic in the GZSD problem. In fact, one key argument of this paper is that traditional detectors/proposal generators often suppress unseen objects as backgrounds, and thus cannot detect unseen objects initially. To mitigate this issue, this paper proposes a novel confidence prediction layer which takes a combination of visual features, semantic features, and spatial locations as input to jointly justify the existence of unseen objects. In addition, unlike that exploit a two-step detection setup, our GZSD detector is built on the one-step YOLO detector, which is fast and scalable to large data size. Last but not the least, this paper adopts a substantially more thorough evaluation setting with multiple splits of both ZSD and GZSD experiments, whereas both and are mainly evaluated on the ZSD test setting (only unseen classes) with only one split. We also noticed that only evaluated on the ILSVRC-2017 which is more constrained where most images contain only one object that needs to be detected.
III Methodology
When developing ZS-YOLO, we deliberately do not require it to identify the test object class names for three reasons: (1) The failure of YOLOv2 on unseen classes is caused by missing detection rather than classification. That is, YOLOv2 tends to relegate unseen classes as backgrounds. Thus, the low recall rates on unseen classes becomes the main issue for ZS-YOLO to address; (2) Once the missing detection problem is solved, the classification task of seen and unseen classes with the extracted bounding boxes has already been extensively explored by ZSR with saturated off-the-shelf solutions ; (3) Different to ZSR where the set of unseen classes are given during training, ZS-YOLO works in a more generalized setting where in completely unknown and unconstrained. In other words, its task is to detect any foreground objects whose classes are not limited by any pre-defined set.
III-B ZS-YOLO Network Architecture
The architecture of the proposed ZS-YOLO is shown in Fig. 2. It consists of four modules: a feature extraction module, an object localization module, a semantic prediction module, and an objectiveness confidence prediction module. Same as YOLOv2 , ZS-YOLO is fully convolutional , i.e. constructed by only convolutional layers and pooling layers.
We take the backbone architecture of YOLOv2 (named as Darknet-19) as the CNN subnet due to its superiority on handling multi-scale objects with a specifically designed passthrough layer applied on fine-grained features. In practice, any other popular CNN model, e.g. ResNet , Inception-V3 , VGG16 , can be replaced here as the base network. Same as , we resize input images into , and our feature extractor outputs a feature tensor, denoted as , in shape of for each input image. After feature extraction, our network is then divided into two branches, which perform object localization and semantic attribute prediction respectively.
III-B2 Object Localization
Similar to , we divide each image into cells, each represented by a -dim vector in . Each cell is assigned with 5 anchor boxes with pre-defined aspect ratios, and needs to predict 4 coordinate offsets for each anchor box. This is implemented by a fully convolutional layer with filters of kernel size. The output tensor, denoted as , is thus in shape of . For locating a predicted bounding box, assuming the cell location is denoted by , the width and height of an anchor box is , and the predicted offsets are , the sigmoid function, then the location of this prediction is calculated by:
III-B3 Semantic Prediction
III-B4 Detection of Overlapping Objects.
Observe that our method allows for detection of overlapping objects either from the same or of different classes. This is because each bounding box outputs its own attribute prediction. Therefore, even if the objects are overlapping, such as the motorbike and person in Fig. 2, the predicted attributes correspond to different bounding boxes and the outputs are different, enabling recognition. Our semantic loss (discussed later) can further enforce diversity among predictions. In addition, even when bounding boxes are in the same location, as a consequence of YOLOv2, they arise from different cells, and so the attributes will be learned from different visual features resulting in different outputs.
III-B5 End-to-End vs. Two-Step Recognition.
Our model could be viewed as an embedding based end-to-end zero-shot localization and recognition method, when semantic attributes for all target classes are available. This follows from the fact that our model produces two outputs at test-time, a bounding box with visual features, and a semantic output. The semantic output is not used whenever target class semantic attributes are unavailable, such as in-the-wild recognition problems. Nevertheless, observe that the semantic output for test examples is always available, and incorporates both test visual features as well as semantic attributes for seen classes. Consequently, the semantic output can be directly compared against ground-truth semantic attributes such as nearest-neighbor or cosine similarity to output a label. In an embedding method, the visual features are mapped to semantic space and these attributes are compared against ground-truth. Viewed in this way, ours is an embedding based method , a standard approach for generalized zero-shot learning. Compared to the cascaded or two-step ZSR methods, the computation of the NN classifier is negligible, and the performance is similar to a two-step approach, as we show in the experiments. Moreover, other loss functions widely used in zero shot recognition methods (e.g. max-margin losses) may also be leveraged for the semantic prediction for our setting.
III-B6 Confidence Prediction
The final component of our network is to predict a confidence score, , associated with each bounding box proposal. Existing detectors, e.g. SSD and YOLOv2 , predict the confidence score directly from the CNN feature map . However, such a strategy produces low confidences for unseen objects visually different to the seen training data, and thus suffering from low recalls. To address this problem, ZS-YOLO predicts the confidence utilizing information from both visual and semantic domain. Additionally, the bounding box coordinates can also be explored as useful sources for predicting the objectiveness confidence, since foreground objects are often located on certain locations (e.g. center rather than corners) of an image.
More specifically, we concatenate all output tensors from previous modules (, , ) as a combined multi-modal input tensor with shape , and forward it into a convolutional layer with 5 single dimensional filters, generating an output tensor , each bit representing the predicted confidence score for the corresponding one of the bounding boxes on an image.
III-B7 Choice of YOLOv2 Detection Architecture.
In our implementation we choose YOLOv2 as the backbone network for three reasons: (a) As a single stage detector, YOLOv2 has a very similar structure with SSD. The only major difference is that YOLOv2 uses Darknet as the feature extractor while SSD uses VGG. As reported in , YOLOv2 reaches higher accuracy on VOC2007 test set while the speed is also faster than SSD512. It is thus not difficult to apply the same idea on SSD. (b) Another popular single stage detector, RetinaNet, mainly leverages focal loss. Our experiments with penalizing focal loss did not demonstrate noticeable benefits on performance, as we will discuss later in experiments. (c) Many two-stage detectors, like Faster-RCNN, are based on the so called Region Proposal Network (RPN), which generates bounding box proposals in the first stage. This two-stage setup makes the detector very slow compared to single stage network. The RPN outputs the bounding boxes and the confidence scores without doing any classification. Thus RPN could be seamlessly integrated into ZS-YOLO by substituting the detector module in ZS-YOLO with an RPN module. Nevertheless, while these offer a number of options, our focus here was to understand and quantify the benefits of semantic attributes for zero-shot detection, and less on comparison of different detectors.
III-C Zero-Shot Detection Losses
With the network architecture defined, we minimize an objective function with a multi-task loss specifically designed for ZSD. Our overall objective loss function is a weighted sum of an object localization loss, a semantic loss and a a confidence loss.
III-C2 Semantic Loss
The semantic loss is designed so that our network could learn a semantic vector representation for each seen class attributes which could be generalized to unseen classes during testing time.
III-C3 Confidence Loss
Finally, the confidence loss is imposed so that our network will predict high objectiveness confidence scores on foreground bounding box proposals and low scores for bounding boxes containing only image backgrounds. Formally our confidence loss is defined as:
Note that we put a minus 0 explicitly in Eq.(3-4) to emphasize the training objective that when the predicted bounding box contains only backgrounds, the max semantic similarity score and the confidence score should be close to zero.
Finally, for each training image, its total objective loss is calculated by a weighted sum of all the above three losses with weights . In our experiments, we set all of them to be 1. Our network is trained by evaluating the empirical loss over the entire training set via stochastic gradient decent.
III-D Training Details
We use the first 16 convolutional layers from Darknet-19 pretrained on ImageNet 1K-class dataset and add 3 randomly initialized convolutional layers as our feature extractor.
Validity of Pretraining Weights for Detection. Note that the pretrained weights correspond to the Darknet weights on ImageNet classification task, and not the detection challenge task. Thus no localization supervision is provided in pretraining, which is the focus of this paper. Also, when training the detector, no labels or bounding boxes for unseen objects are accessible. Therefore, observe that the visual features for unseen classes are viewed as background unless they are learned from seen classes. Consequently, the representation embedded in the pretrained weights will tend to be suppressed due to lack of labels. In addition, in the experiments, both YOLOv2 and ZS-YOLO are initialized with pretrained Darknet weights for fair comparison. We also point out that our focus here is to understand benefits of semantic attributes in the best situation, where other components, modules and weights are suitably well-chosen. Indeed, the general problem of jointly optimizing the entire system is important but out of scope of this paper. Based on what we observe the issue of the ImageNet pretrained weights does not appear to be a dominant aspect of the large performance gap between seen and generalized zero-shot detection.
The activation functions of the semantic predictor and confidence predictor are linear and all other layers use a leaky rectified linear activation function :
The network is trained end-to-end for 420 epochs on the training splits from PASCAL VOC 2007 and 2012 and MS COCO. We use a SGD optimizer with the batch size 64, momentum 0.9 and a decay of 0.0005. As for the learning rate, in the first 5 epochs, we set it to 0.0001 because the model often diverges if it starts with a high learning rate. Then we increase it to 0.001 and train the model for 195 epochs. Afterwards we decrease it to 0.0001 for 110 epochs, and train it with 0.00001 for the final 110 epochs.
IV Experiments
The average precision is then defined as the mean precision at a set of eleven equally spaced recall levels :
Unlike standard mAP used in Pascal VOC, we measure AP over all classes since we do not classify the objects. There is no way to get a class-specific precision and so class-level AP and mAP are not computed. The AP reflects the overall detection performance on all possible objects. Nevertheless, as we have already mentioned in Sec. III-B5, we could conceivably use our system as an embedded based recognition method, whenever semantic attributes for ground-truth are provided. With this in mind we also tabulate recognition results for the sake of completeness.
Nevertheless, we specifically care about AP of unseen classes since our goal is to leverage YOLOv2 to detect more unseen objects. We also report the average F-score as an auxiliary measurement since it reflects the average performance of the detector at different confidence thresholds The average F-score is computed by averaging the F-scores over a set of 101 equally spaced confidence thresholds , at each confidence threshold, the F-score is defined as:
The average F-score can reflect the robustness of overall performance for the detector.
To evaluate the zero-shot detection performance of the proposed ZS-YOLO model, we chose PASCAL VOC 2007 and 2012 datasets and Microsoft COCO object detection dataset due to their popularity among object detection literature.
This dataset contains 20 object classes in total, and each class is labeled with -dimensional semantic attributes published by . The binary instance level attributes are provided in and we compute the class-level attributes by averaging over all instances in the class. There is no standard seen/unseen split on PASCAL VOC for object detection, we thus built our own splits. Since PASCAL VOC is a relatively small dataset we utilize PASCAL primarily for several ablative studies. In our first experiment our goal is to quantify the impact of increasing unseen classes. We discuss other ablative studies in Sec. IV-C. For this study, we held out different numbers of unseen classes and utilized the rest for training. These unseen classes were selected based on their diversity in class types (e.g. avoiding similar classes such as dog and cat from being both included as unseen). We thus ended up with three different seen/unseen class balances: respectively 15/5, 10/10, and 5/15. During model training for each split, we collect all the images which only contain seen object classes in the train/val partition of PASCAL VOC 2007 and 2012 dataset as training data. For testing, we use three different data configurations, named as Test-Seen, Test-Unseen and Test-Mix. Test-Seen data is constructed by the images from VOC2007 test partition which only contain seen classes; Test-Unseen data contains all the images having only unseen class objects from both train/val/test partitions of VOC2007 and train/val partition of VOC2012; All the other images, which contain both seen and unseen objects, go into our Test-Mix data. Test-Unseen corresponds to the standard zero shot detection task, where only unseen classes are in the data. The ability of discovering unseen objects can be evaluated on Test-Unseen. Test-Mix is a more sophisticated situation where both seen and unseen objects appear in the image. The detector needs to identify unseen objects (assign high confidence score) even in the presence of seen objects that the detector has evidently been optimized to detect during training. The high score of seen objects may suppress the prediction on unseen objects resulting in poor unseen scores. Finally, Test-Seen is a conventional detection dataset and quantifies the ability to ensure good detection for seen (supervised) classes. We list the components of our dataset in TABLE I in more detail.
Class-level Attributes vs. Object-level Attributes: There are three reasons for us to adopt the averaged class-level attribute representation, instead of the original object-level attributes by . (1) During the training stage of GZSD/R models, while the class-level averaging might result in reducing the distinctiveness for each object, the annotation noises and variations also decrease after averaging. As observed in our experiments, adopting the class-level attribute representation as ground-truth label makes the training procedure more stable and results in better GZSD performance (similar observations also made by previous ZSR works, e.g.). This might be easy to understand since the task is to achieve a good GZSD/R performance globally, rather than to distinguish each instance within a specific class. (2) During the testing stage of GZSD/R models, only the class-level semantic information (attributes) of the unseen classes is provided, and the object-level unseen attributes are impossible to acquire beforehand. (3) Finally, the aPY dataset only labels a fraction of PASCAL VOC objects and there are still a large amount of objects in PASCAL VOC without any object-level labels.
IV-A2 MS COCO Dataset
MS COCO dataset has 80 classes which include all 20 classes in PASCAL VOC. While an attribute dataset for MS COCO has been published in , we could not utilize it for two reasons: (1) The dataset only labels 29 out of the 80 object classes. The number is too limited to conduct extensive ZSD experiments on COCO. (2) Many attributes provided in are not visually meaningful, e.g., professional,useful, friendly, functional, and thus not suitable for visual detection/recognition. An alternative to attributes, word2vec (w2v), is noisy and performs poorly. To extend our model to train on more classes, we propose to learn transformation , that maps w2v features (300 dimensions) onto a lower-dimensional w2vR space (25 dimensions, i.e. ), which is constrained to mirror VOC attribute similarity on the 20 common classes between MS COCO and PASCAL VOC. Specifically, for class attribute , and w2v vector , we seek such that
This problem can be solved by using any one of the rank-approximate methods. More detailed evaluation can be found in TABLE V, Section IV-C.
MS COCO has more classes, while PASCAL VOC has few classes and this offers several options for conducting ablative studies. First, with PASCAL VOC we can tabulate effect of different ratios of seen/unseen classes. For the largest ratio, we hope to see less difference between ZS-YOLO and YOLOv2 (for test-mix or test-seen) because visual features are significantly stronger than semantic information. In contrast for very small ratio there are more unseen classes and very few seen classes and so we hope to see generally poor performance. Second, for MS COCO, since there are a large number of classes we can conduct a different type of experiment, namely, how unseen performance varies as we see more seen classes. For this reason, we tabulate performance for a fixed set of unseen classes as a function of more seen classes available for training. Third, MS COCO also allows for us to study the difference between YOLOv2 and ZS-YOLO at a larger scale in the presence of significantly more training data.
Based on this motivation, we first manually selected 20 classes that were sufficiently diverse for unseen categories. We then chose training categories that were semantically most similar in word2vec representation to the unseen categories. The number was increased in increments of 20. Similar to PASCAL VOC, we collected all images which only contain seen object classes in the train partition of MS COCO dataset as training data. Accordingly, Test-Seen is constructed by images only containing seen objects in validation partition of MS COCO; Test-Unseen contains all images only containing unseen objects in train/val partitions of MS COCO and all the images containing both seen and unseen objects in MS COCO train/val dataset, go into Test-Mix data. The components of the dataset are detailed in TABLE I.
IV-B Zero-Shot Detection Evaluation
We first evaluate the proposed ZS-YOLO model by comparisons on all seen/unseen splits as well as Test-Seen/Unseen/Mix configurations. As a baseline comparison, we also train a YOLOv2 as a standard fully-supervised detector (i.e. with both class and bounding box labels) using the same training splits. During test time, since we measure the average precision without classification, the classifier module in YOLOv2 is detached and the bounding box predictions are made based on confidence score.Our comparative results are shown in TABLE II as well as Fig. 3.
TABLE II evidently shows the advantage of the proposed ZS-YOLO model, especially on detecting unseen object classes. Compared with YOLOv2, our model has a higher AP on Test-Unseen data on all the three different seen/unseen splits, e.g. we improve from YOLOv2’s to on PASCAL VOC (10/10 split), and similarly, from to on MS COCO (60/20 split). Our hypothesis is that the main reason for this performance gain in unseen AP is that ZS-YOLO predicts objectiveness confidence score based on both visual as well as semantic information, and thus effectively avoids suppressing unseen objects with novel visual features but closely-related semantic meanings. Observe that YOLOv2’s performance is uneven as we increase seen class categories. For instance, on MS COCO dataset, YOLOv2 achieves the best performance on Test-Unseen () when trained on 40 seen classes, but worse when trained with 20 and 60 classes ( and respectively). We attribute this to the fact to an under and over utilization of seen classes. Namely, in presence of few seen categories, without semantic knowledge, it is hard to generalize to unseen classes. On the other hand, in the presence of large seen training classes, YOLOv2 learns to reject unseen classes well and tends to classify unseen class as background. On the other hand, since ZS-YOLO exploits semantic feature to overcome this issue, we observe a continuous improvements of ZS-YOLO on Test-Unseen as the number of training seen classes increases (from to on MS COCO). Additionally, we observe on Test-Seen partition that both models suffer from a performance drop when seen classes increase from 40 to 60. This is because many difficult classes are added (e.g. stop sign, remote, etc) which impede both models and lowers the average precision.
From Fig. 3, observe that ZS-YOLO’s recall on unseen objects is substantially larger than YOLOv2 on both datasets thanks to its semantic attributes based detection framework. We illustrate this with the PASCAL VOC 5/15 split. Fig. 3 shows that at a confidence threshold of , ZS-YOLO can still detect about of the total unseen objects, whilst original YOLOv2’s recall rate is only . While ZS-YOLO achieves higher AP on unseen data with significantly improved recall rates, with degradation in precision on seen data. For instance, on PASCAL VOC ZS-YOLO loses (10/10), (15/5) and (5/15) AP on Test-Seen compared to YOLOv2. We posit that this is fundamental. Indeed, in order to detect more unseen data, we must inevitably accept more background bounding boxes to improve unseen object detection. However, if both models are evaluated with the Test-Mix data with both seen and unseen classes, we found that in general ZS-YOLO achieves better performance compared to YOLOv2.
We can observe from Fig. 3 that except for VOC 5/15 split, ZS-YOLO has a higher precision than YOLOv2 in any level of recall. The 5/15 split is a special case where the number of seen attributes are somewhat insufficient to generalize to unseen objects. For example, on VOC 10/10 split, at the recall level 60%, YOLOv2 has precision 50% while ZS-YOLO is 61%. Also, ZS-YOLO has higher recall than YOLO on any level of confidence score. Although YOLOv2 achieves the same recall as ZS-YOLO at a smaller confidence threshold, it suffers from loss on precision. Consequently, we do benefit from semantic predictions, specifically, when we have sufficient number of training classes and, as noted later, when the semantic unseen and seen attributes are related Table VI. Noticeably, ZS-YOLO is robust, with a high recall in a wide range of confidence, meaning its performance does not degrade for a wide range of confidence scores. This can also be justified by the average F-score in Table. II.
It might be surprising that the performance gap is not large over YOLOv2 on PASCAL VOC as revealed in TABLE II and Fig. 3. Importantly, we argue that this is primarily because other than the unseen object classes among the 20 annotated ones we artificially held out during training, there are many more unannotated object classes in the PASCAL VOC dataset which do not have their corresponding ground truth boxes. Thus, they are also counted as false positives even though ZS-YOLO does successfully detect them as foregrounds. Unfortunately, these cases cannot be quantitatively measured. We thus show in Fig. 4 extensive qualitative results, where ZS-YOLO not only succeeds in detecting artificial unseen object classes, but also those that are truly unseen objects that have no annotation. These successful detections will be measured as false positives in TABLE II and Fig. 3, resulting in lower AP. Therefore, ZS-YOLO’s performance gain can be expected to be larger than YOLOv2 in practice.
IV-C Ablative Study on PASCAL VOC
To further evaluate the effect of the semantic prediction side task during model training and its impact on the learnt visual features, we conducted ablative analysis on the proposed ZS-YOLO model. Specifically, we trained two different models for the PASCAL VOC 10/10 split: (1) ZS-YOLO (visual): Instead of using the combined multi-model input (see Fig. 2) for final confidence score prediction, we removed the semantic prediction module completely from the network architecture, so the final confidence prediction is based purely on without semantic information; and (2) ZS-YOLO (semantic): In this model, we remove visual features from confidence prediction input by only feeding into the confidence prediction module. Other than the above two models, we consider original YOLOv2 which only uses for confidence prediction as a third baseline. The comparative results are shown in TABLE III.
From TABLE III we observe that utilizing both visual features as well as semantic attributes as multi-modal clues improves the detection performances on unseen data, compared to detection models that exclusively use single-modal information to infer objectness confidence. Specifically, on unseen data the ZS-YOLO(full) model that explores both visual and semantic domains achieves AP on Test-Unseen data, which is the highest among all three competitors. By comparing YOLOv2 with ZS-YOLO(visual), we can see that using the bounding box locations as a side information also helps to detect unseen objects, although such information might be redundant for detecting seen objects especially when the visual features are uniquely fine-tuned to those classes. Finally, we can see when we exclusively utilize only semantic information, ZS-YOLO(semantic) generates better detection performance compared to YOLOv2 on unseen data, which again validates the importance of utilizing semantic information when moving toward zero-shot detection. This effect can be visualized in Fig. 5 with t-SNE embeddings of all the bounding box proposals generated on the PASCAL VOC Test-Mix data by using the learned YOLOv2’s visual feature as well as ZS-YOLO’s semantic feature (i.e. ).
IV-C2 Choices of Semantic Prototypes: One-hot & Random
We are also interested in the problem of whether using attributes and word2vec (w2v) as semantic prototypes are indeed superior to other choices. Two semantic prototype spaces were tested for training in this context: (1) One-hot encoding: we use the one-hot class label vectors as semantic prototypes, so that all the prototypes are orthogonal to each other. Note that one-hot vectors were used only to encode seen classes. Since, we disconnect semantic attribute predictor at test-time, the issue of unseen semantic representation does not arise for our situation. Indeed, were we to look at semantic predictions, it is quite likely that at test-time the unseen classes are possibly represented by a combination of seen (one-hot) classes, analogous to the semantic attributes. (2) Random encoding: We generate the prototypes for each class with same dimension as attribute vectors () by randomly sampling from uniform distribution. We run the experiments on the PASCAL VOC 10/10 split and the comparative results are reported in TABLE IV. The results clearly show that by using meaningful semantic prototypes such as attributes and word2vec, the detection performance on unseen objects can be improved. Random encoding prototypes, which contain no semantic information, make the AP of our model 49.0%, significantly worse than original YOLOv2. Meanwhile, by utilizing one-hot encoding semantic prototypes, class level information is introduced and our model can get better detection performance on unseen data.
IV-C3 Effect of Semantic Dimensionality Reduction
To verify our semantic dimensionality reduction method as discussed in Section IV-A, we compared models trained in the reduced word2vec space (w2vR) with the original word2vec space (w2v) on both PASCAL VOC (10/10 split) and MS COCO (20/20 split) datasets. Since the original w2v space (300 dimensions) is highly noisy, which is irrelevant to the visual space, learning such information has the effect of poor convergence behavior for a convolutional neural network. On the contrary, the reduced word2vec space w2vR is defined to preserve similarity structures originally defined in the attribute space which is more visually related. It is evident from TABLE V that models trained on our mapped w2v, w2vR, leads to better performance on detecting unseen objects than the original w2v. In particular, on MS COCO dataset, the model trained on w2vR gets an AP of 40.6%, even higher than model trained on VOC attribute, 38.4%.
IV-C4 Effect of Seen/Unseen Correlations
In zero shot recognition, the unseen and seen classes have to share some common visual appearance such that the classifier can generalize the representation learned from seen classes to unseen. If the unseen classes have low correlation with seen classes, the zero shot task tends to be more difficult. The same phenomenon may also exist in zero shot detection. Therefore, we seek to test the effect of correlation between the attributes of seen and unseen classes on detection performance of unseen objects. To measure this correlation, we define an energy score function of a class split as: .
If the energy score is higher, the unseen classes have higher correlation with the seen classes. To further analyze its effect, we construct two more 10/10 splits (named as 10/10-2 and 10/10-3) and one 5/15 (5/15-2) split experiments over PASCAL VOC. The splits are constructed based on the energy score. 10/10-2 has an intermediate energy score over all 10/10 splits while 10/10-3 has the smallest. 5/15-2 has the smallest energy score and since it is very close to 5/15-1 which already has the highest energy score we don’t need to construct one more. When the energy score is high, visually similar classes seem to appear in seen/unseen split separately, and when it is low, they seem both in seen or unseen. For example, in 10/10-1 split, for some similar class pairs like motorbike-bicycle, bicycle is in seen and motorbike is in unseen. While in 10/10-3 split, they are both in seen. The detailed data splits are available at the first author’s github page. Our results are reported in TABLE VI, with each split’s seen/unseen correlation calculated by score defined in Section IV-A. TABLE VI is consistent with our intuition, that in general when seen and unseen data are more semantically correlated, our ZS-YOLO performs better on detecting unseen classes. For instance, in split 10/10-1 where , ZS-YOLO achieves a AP on Test-Unseen; Whilst on split 10/10-3 where seen/unseen classes are less semantically related (), ZS-YOLO achieves a inferior AP of .
IV-C5 Other Losses
We tried other losses such as the focal loss that has been suggested in the literature to improve detection through hard positive/negative mining. However, we found that while seen object detection improved, we saw a larger loss for unseen objects.
IV-D Semantic Output for Zero Shot Recognition
For gZSL setting, the mean accuracy (1) and mAP in seen and unseen classes are also listed. We emphasize the fact that there are two aspects that contribute to the error. First, error in recognition even when bounding boxes are provided. Second, when bounding boxes are not provided and so both bounding box and recognition is required. Note that GZSL accuracy (namely only the recognition task) for PASCAL VOC is notoriously hard with poor accuracy. Consequently, we have a poor baseline to start with and in addition must also perform detection. Our goal in this context is to quantify error increase suffered by virtue of detection. In our ablative study we find that detection error is a smaller component relative to Generalized recognition error for PASCAL VOC.
By leveraging ZS-YOLO’s semantic prediction, ZS-YOLO + NN reaches the highest AP on almost every class in Test-Unsee, and most classes in Test-Mix. Using the same classifier, ZS-YOLO + reaches higher mAP than YOLOv2. This could be attributed to the ability of detecting unseen objects of ZS-YOLO. In the Test-Mix, ZS-YOLO outperforms YOLOv2 on all unseen classes except for cat. It is worth noting that YOLOv2 + achieves 17.66 on bike but 0 on motorbike, suggesting YOLOv2 suppresses unseen objects even they have visually similar seen classes.
First from Table. VII note that the GZSL error for when bounding boxes are provided is already quite low. While this is not state-of-art, performance on aPY (which is the recognition task on PASCAL VOC) is not much larger with other methods. Our purpose here is less about choosing the best ZSR model and more about detection of unseen objects and so we did not consider other approaches here. Next we observe that both ZS-Y + NN and ZS-Y + achieves much higher mAP on unseen classes. The unseen mAP for ZS-Y + is 6.92%, which is better than 4.33% from ZS-Y + NN. However, we notice that ZS-Y + NN outperforms on 7 unseen classes while ZS-Y + only reaches much higher AP on boat and car. Therefore the attribute prediction in ZS-Y is quite accurate and it outperforms cascaded ZSR. ZS-Y + NN outperforms YOLOv2 + on both seen and unseen classes. The error of full zero shot detection comes from the combination of erroneous localization plus misclassification, as the accuracies from are very low for some classes especially in gZSL setting, even though it has no localization error. The results clearly show that our ZS-YOLO can be combined with zero shot recognition methods easily to construct a full zero shot detector and outperforms the original YOLOv2 in the sense of detecting unseen objects.
Discussion: We focused on detection aspect of GZSD. In this context, we adopted a simple nearest neighbour model for classification. Nevertheless, we agree that the performance of GZSR classifiers also contributes to the final GZSD performance. Specifically, a key challenge in GZSR is the gap between the semantic domain and the visual domain. There already exists various efforts to bridge such a gap, e.g. . Recently, we have found that a promising solution to mitigate the divergence between semantic and visual domains might be to learn a low-dimension ’visually semantic’ embedding that quantifies existence of a prototypical part-type in the presented instance of seen classes. As a possible future direction toward better GZSD performances, one could combine the proposed ZS-YOLO detector with the more powerful ZSR classifiers described above instead of the nearest neighbor classifier adopted in this paper.
V Conclusion
We proposed a novel Zero-Shot YOLO method for unseen object detection that retains the principle advantages of YOLOv2 ’s efficiency and performance for seen object detection and extends it non-trivially for unseen object detection. While YOLOv2 is a state-of-art object detection algorithm, its effectiveness hinges on having access to fully annotated datasets, which is unrealistic to expect as as we move towards large-scale object detection. YOLOv2’s effectiveness in learning sharp visual features for accurately detecting objects that were seen during training negates its advantages for unseen object detection. To overcome these drawbacks, we propose to build upon YOLOv2’s network architecture through seamless fusion of semantic information with the visual domain so that, semantically similar object attributes are also reflected in the learned visual features. Empirically we demonstrated improved performance on two datasets.
Acknowledgment
The last author would like to thank Dr. Ziming Zhang for initial helpful discussions on zero shot detection and Dr. Tolga Bolubasi for preliminary paper draft and discussions on experiments. The authors would like to thank the Associate Editor and the reviewers for their constructive comments. This work was supported by the Office of Naval Research Grant N0014-18-1-2257, NGA-NURI HM1582-09-1-0037 and the U.S. Department of Homeland Security, Science and Technology Directorate, Office of University Programs, under Grant 2013-ST-061-ED0001.