Object Detection in Videos by High Quality Object Linking

Peng Tang, Chunyu Wang, Xinggang Wang, Wenyu Liu, Wenjun Zeng, Jingdong Wang

Introduction

Detecting objects in static images has achieved significant progress due to the emergence of deep convolutional neural networks (CNNs) . However, object detection in videos brings additional challenges due to degraded image qualities, e.g. motion blur and video defocus, leading to unstable classifications for the same object across video. Therefore, many research efforts have been allocated to video object detection by exploiting temporal contexts , especially after the introduction of the ImageNet video object detection (VID) challenge.

Many previous methods exploit temporal contexts by linking the same object across video to form tubelets and aggregating classification scores in the tubelets . They first use static image detectors to detect objects in each frame, and then link these detected objects by checking object boxes between neighboring frames, according to the spatial overlap between object boxes in different frames or predicting object movements between neighboring frames . Very promising results are obtained by these methods.

However, the same object changes its locations and appearances in neighboring frames due to object motion, which may make the spatial overlap between boxes of the same object in neighboring frames not sufficient enough or the predicted object movements not accurate enough. This influences the quality of object linking, especially for fast moving objects. By contrast, in the same frame, it is obvious that two boxes correspond to the same object if they have sufficient spatial overlaps. Inspired by these facts, we propose to link objects in the same frame instead of neighboring frames for high quality object linking.

In our method, a long video is first divided into some temporally-overlapping short video segments. For each short video segment, we extract a set of cuboid proposals, i.e. spatio-temporal candidate cuboids which bound the movement of objects, by extending the region proposal network for static images to a cuboid proposal network for short video segments. The objects across frames lying in a cuboid are regarded as the same object. The main benefit from cuboid proposal is to enable object linking in the same frame and it alone yields minor detection performance improvement.

For each cuboid proposal, we adapt the Fast R-CNN to detect short tubelets. More precisely, we compute the precise box locations and classification scores for each frame separately, forming a short tubelet representing the linked object boxes in the short video segment. We compute the classification score of the tubelet, by aggregating the classification scores of the boxes across frames. In addition, to remove spatially redundant short tubelets, we extend the standard non-maximum suppression (NMS) with a tubelet overlap measurement, which prevents tubelets from breaking that may happen in frame-wise NMS. Considering short range temporal contexts by short tubelets benefits detection, see Fig. 1 (b).

Finally, we link the short tubelets with sufficient overlap across temporally-overlapping short video segments. If two boxes, which are from the temporally-overlapping frame (i.e. the same frame) of two neighboring short tubelets, have sufficient spatial overlap, the two corresponding short tubelets are linked together and merged. We exploit the object linking to improve the classification quality by boosting the classification scores for positive detections through aggregating the classification scores of the linked tubelets. As shown in Fig. 1 (c), the detection results can be further improved by considering long range temporal contexts.

Elaborate experiments are conducted on the ImageNet VID dataset . Our method obtains mAP 74.5%74.5\% training on the VID and 80.6%80.6\% training on the mixture of VID and DET. The results outperform both the static image detector and the previous best performed methods. In particular, our method obtains 8.8%8.8\% absolute improvement compared with the static image detector for fast moving objects.

Related Work

The task of object detection in both images and videos has been widely studied in the literature. We mainly review related works on video object detection and classify them into three categories by how they use the temporal contexts.

Feature Propagation w/o Object Linking. In , the features of the current frame are augmented by aggregating features propagated from neighboring frames. The methods in use the optical flows to spatially align features in different frames for feature propagation. Bertasius et al. propagate features by using the deformable convolutional network across space and time. Xiao and Lee adapts the Conv-GRU to propagate features from neighboring frames. Feature propagation is also exploited in to speed up the object detection. The authors propose to compute the feature maps (using a very deep network with high computation cost) for the key frames and propagate the features to non-key frames by computing the optical flows using a shallow network which takes less time. These methods are different from ours because they do not perform object linking.

Feature Propagation w/ Object Linking. The tubelet proposal network computes tubelets by first generating static object proposals in the first frame and then predicting their relative movements in following frames. The features of the boxes in the tubelets are propagated to each box for classification by using a CNN-LSTM network. Apart from the feature propagation w/o object linking, Wang et al. also link object in neighboring frames for feature propagation. More precisely, the relative movements in neighboring frames are predicted for each proposal in the current frame, and the features of the boxes in neighboring frames are propagated to the corresponding box in the current frame by average pooling. Unlike these methods, we link objects in the same frame and propagate box scores instead of features across frames. Besides, we directly generate the spatio-temporal cuboid proposals for video segments rather than per-frame proposals in .

Score Propagation w/ Object Linking. The method in proposes two kinds of object linking. The first one tracks the detected box in current frame to its neighboring frames to augment their original detections for higher object recall. The scores are also propagated to improve classification accuracy. The linking is based on the mean optical flow vector within boxes. The second one links objects into long tubelets using the tracking algorithm and then adopts a classifier to aggregate the detection scores in the tubelets. The Seq-NMS method links objects by checking the spatial overlap between boxes in neighboring frames without considering the motion information and then aggregates the scores of the linked objects for the final score. The method in simultaneously predicts the object locations in two frames and also the object movements from the preceding frame to the current frame. Then they use the movements to link the detected objects into tubelets. The object detection scores in the same tubelet is reweighed by aggregating the scores in some manner from the scores in that tubelet.

Our method belongs to the third category. The main contribution of our work is that we link objects in the same frame instead of neighboring frames in previous methods . In addition, to achieve our goal, we develop a series of methods such as cuboid proposal network which have not been explored in previous methods.

Method

The task of video object detection is to infer the locations and classes of the objects in each frame of a video {I1,I2,…,IN}\{\mathbf{I}^{1},\mathbf{I}^{2},\dots,\mathbf{I}^{N}\}. To obtain high quality object linking, our method proposes to link objects in the same frame, which can be used to improve the classification accuracy.

Given a video divided into a series of temporally-overlapping short video segments as the input, our method consists of three stages: (1) Cuboid proposal generation for a short video segment. This stage aims to generate a set of cuboids (containers) which bound the same object across frames as shown in Fig. 2. See Section 3.1. (2) Short tubelet detection for a short video segment. For each cuboid proposal, the goal is to regress and classify a short tubelet which is a sequence of bounding boxes with each box localizing the object in one frame. The spatially-overlapping short tubelets are removed by tubelet non-maximum suppression. The short tubelet is a representation for linked objects across frames in a short video segment, as illustrated in Fig. 2. See Section 3.2. (3) Short tubelet linking for the whole video. This stage, depicted in Fig. 3, links the temporally-overlapping short tubelets to link objects in the whole video, and refines the classification scores of the linked tubelets. See Section 3.3. The first two stages, cuboid proposal generation and short tubelet detection, generate temporally-overlapping short tubelets, and thus ensure that we can link objects in the temporally-overlapping frame (i.e. the same frame) in the short tubelet linking stage.

We modify the region proposal network (RPN) method in Faster R-CNN and introduce the cuboid proposal network (CPN) method for computing cuboid proposals. Unlike the conventional RPN where the input is usually a single image, our method takes the KK frames as the input to the CPN. The output is a set of whkwhk cuboid proposals, regressed from a w×hw\times h spatial grid, where there are kk reference boxes at each location, and each cuboid proposal is associated with an objectness score.

2 Short Tubelet Detection

We use the 2D2D form of the cuboid proposal, as the 2D2D box (region) proposal for each frame in this segment, which is classified and refined for each frame separately.

Considering a frame Iτ\mathbf{I}^{\tau} in this segment, we follow Fast R-CNN to refine the box and compute the classification score. We start with a RoI pooling operation, where the input is a 2D2D region proposal b\mathbf{b} and the response map of Iτ\mathbf{I}^{\tau} obtained through a CNN. The RoI pooling result is fed into a classification layer, outputting a {C+1}\{C+1\}-dimensional classification score vector yτ\mathbf{y}^{\tau}, where CC is the number of categories and 11 corresponds to the background, as well as a regression layer, from which the refined box is obtained.

The resulting KK refined boxes for the KK frames form the short tubelet detection result over this segment, T=(bt,bt+1,…,bt+K−1)\mathcal{T}=(\mathbf{b}^{t},\mathbf{b}^{t+1},\dots,\mathbf{b}^{t+K-1}). The classification score of this tubelet is an aggregation of the scores over all the frames,

where aggregation⁡(⋅)\operatorname{aggregation}(\cdot) could be a mean⁡\operatorname{mean} operation. We empirically find that Aggregation⁡(⋅)=12(mean⁡(⋅)+max⁡(⋅))\operatorname{Aggregation}(\cdot)=\frac{1}{2}(\operatorname{mean}(\cdot)+\operatorname{max}(\cdot)) performs the best.

To remove spatial redundant short tubelets, we extend the standard non-maximum suppression (NMS) algorithm to a tubelet NMS (T-NMS) algorithm to remove spatially-overlapping short tubelets in the same segment. This strategy prevents tubelets from breaking by frame-wise NMS which removes 22D boxes for each frame independently. The main point lies in how to measure the spatial overlap between two tubelets. We define it on the base of the overlap between the boxes in the same frame. Given two tubelets, Ti=(bit,bit+1,…,bit+K−1)\mathcal{T}_{i}=(\mathbf{b}^{t}_{i},\mathbf{b}^{t+1}_{i},\dots,\mathbf{b}^{t+K-1}_{i}) and Tj=(bjt,bjt+1,…,bjt+K−1)\mathcal{T}_{j}=(\mathbf{b}^{t}_{j},\mathbf{b}^{t+1}_{j},\dots,\mathbf{b}^{t+K-1}_{j}), the spatial overlap is computed as

where IoU⁡(biτ,bjτ)\operatorname{IoU}(\mathbf{b}^{\tau}_{i},\mathbf{b}^{\tau}_{j}) is the intersection over union between biτ\mathbf{b}^{\tau}_{i} and bjτ\mathbf{b}^{\tau}_{j} for frame τ\tau. We choose this measurement because two short tubelets are not the same even if only one pair of corresponding boxes do not have sufficient overlap.

3 Short Tubelet Linking

Our method divides a video into a series of temporally-overlapping short video segments of length KK with stride K−1K-1:

Considering two temporally-overlapping short tubelets: the iith tubelet from the mmth segment and the i′i^{\prime}th tubelet from the (m+1)(m+1)th segment:

we link them if the spatial overlap between bitm+K−1\mathbf{b}_{i}^{t_{m}+K-1} and bi′tm+K−1\mathbf{b}_{i^{\prime}}^{t_{m}+K-1} from the temporally-overlapping frame (i.e. the same frame) is larger than a pre-defined threshold.

We perform a greedy short tubelet linking algorithm. Initially, we put the short tubelets from all short video segments into a pool and record the corresponding segment for each tubelet. Our algorithm pops out the short tubelet T\mathcal{T} with the highest classification score from the pool. We check the IoU of the boxes over the temporally-overlapping frame between T\mathcal{T} and its temporally-overlapping short tubelets. If the IoU is larger than a threshold, fixed as 0.40.4 in our implementation, we merge the two short tubelets into a single longer tubelet, remove the box with the lower score for the overlapping frame, update the classification score for the merged tubelet according to Eq. (2) for better classification, and record the corresponding segment (a combination of the corresponding two video segments). We then push the merged tubelet into the pool. This process is repeated until no more tubelets can be merged. Fig. 3 gives the examples of linking short tubelets to form long tubelets.

The tubelets remaining in the pool form the video object detection results: the score of the tubelet is assigned to each box in the tubelet, and the boxes from all the tubelets associated with a frame are regarded as the final detection boxes for the corresponding frame.

4 Implementation Details

Cuboid Proposal. The base network is ResNet-101101 pre-trained on the ImageNet classification dataset : we remove all layers after the Res5c layer and replace the convolutional layers in the fifth block by dilated ones to reduce the stride from 3232 to 1616. On the basis of the base network, we add a convolutional layer with 512512 filters of 3×33\times 3, and use two convolutional layers of 1×11\times 1 to regress the offsets and predict the objectness scores for cuboid proposals. The network is split into two sub-networks: the first one has two residual blocks pass each frame separately to obtain frame-specific features which are concatenated as input of the second sub-network with three residual blocks.

We use four anchor scales 64264^{2}, 1282128^{2}, 2562256^{2}, and 5122512^{2} with three aspect ratios 11:11, 11:22, and 22:11, resulting in 1212 anchors at each location in total. The length KK of each video segment will be studied in our experiments. The loss function is the same as that in the standard RPN : the cross-entropy loss for classification and the smoothed L1 loss for regression. The training targets are the ground truth cuboids as defined in Eq. (1). The NMS threshold 0.70.7 is chosen and at most 300300 proposals are kept for the detection network training/testing. In the testing stage, if the number of frames in the last segment is smaller than KK, we pad the segment by some frames copied from the last frame.

Short Tubelet Detection. The base network is the same as it for cuboid proposal. We use RoI pooling to extract 7×77\times 7 response maps from the layer Res5c, followed by two fully-connected + ReLU layers (10241024 neurons). We use one fully-connected layer for classification and another fully-connected layer for bounding box regression. Following the Fast R-CNN , we train the network with online hard example mining . The difference between our short tubelet detection training and the Fast R-CNN training is in the ground truth matching. In particular, we match a cuboid proposal to a ground truth box if the IoU between the cuboid proposal and a ground truth cuboid is larger than a threshold (typically 0.50.5). This is because the CPN is trained for cuboids, which makes cuboid proposals hard to match ground truth boxes directly. This matching strategy also ensures that a cuboid proposal corresponds to the same object in different frames. The training targets are still ground truth boxes (rather than ground truth cuboids) because we want to get accurate object locations in each frame. During testing, the T-NMS threshold is set to 0.40.4.

Training. We use SGD to train the cuboid proposal network and the short tubelet detection network. We initialize the weights of the newly added layers by a zero-mean Gaussian distribution whose std is 0.010.01. Images are resized to shorter side 600600 pixels for both training and testing. We set the mini-batch size to 88, the learning rate to 1×10−31\times 10^{-3} for the first 4040K iterations and 1×10−41\times 10^{-4} for the next 2020K iterations, and the momentum to 0.90.9. We do not find the gain from sharing the base networks for the cuboid proposal network and the short tubelet detection network, so we simply train them separately. Our implementation is based on the Caffe deep learning framework on a TitanX (Pascal) GPU.

5 Discussions

Action Detection. The tasks of spatio-temporal action detection and object detection in videos are similar to some extent. The purpose of spatio-temporal action detection is to localize and classify actions in each video frame. Some solutions to action becomes similar to video object detection and some of them can also be cast into the object/action linking framework. For instance, linking through neighboring frames, which is studied in video object detection , is also explored in . We find that only the contemporary work in action detection adopts the scheme of linking through temporally-overlapping frames, and its short tubelet detection scheme, similar to , is different from our cuboid proposal based method. It should be noted that although the solution frameworks of the two problems are similar in high level, the research focuses are different: action detection is more about capturing the motion from the temporal signals and understanding an action from a single frame can be ambiguous (e.g. sitting down or standing up) , whereas video object detection can be done in a single frame and the temporal information is introduced to improve results in some frames of degraded image qualities. As validated in the later empirical results, the state-of-art action detection method , whose framework is similar to our method, performs poor in video object detection.

Multi-object Issue. It is possible that one cuboid contains multiple objects, because a cuboid tends to occupy a larger region than the object. However, as observed in our experiments, this problem has almost negligible influence on the detection performance. This is because our overlapping-based short tubelet linking can be accomplished when a video segment only has two frames. In this case, each cuboid, in most cases, contains only one object. It is worth noting that it is not necessary to use video segments longer than two frames, because short two-frame segments already support overlapping-based short tubelet linking.

Boundary Issue. It is possible that in some frames an object may appear or disappear. As a result, the boundary issue occurs in the short tubelet detection stage. More precisely, in short tubelet detection, KK successive frames share the same proposals and proposal classification scores. Take K=2K=2 as an example, there are two boundary frames, i.e. the one frame before an object appears and the one frame after the object disappears. The proposal classification scores of the two boundary frames will be enhanced according to Eq. (2), which will result in false positives in these two frames. This problem does not occur in the short tubelet linking stage because we allow broken links for long range linking. Actually in real applications, the sequence where the object continuously appears is not short in most cases, and then the two boundary frames will not affect the performance that much. We investigate the VID dataset and find that this boundary issue only leads to small performance drop (up to 0.57%0.57\%).

Experiments

We use the ImageNet VID dataset which was introduced in the ILSVRC 2015 challenge. The dataset contains 3030 object classes which cover different movement types and different levels of clutterness. The dataset has 53545354 videos which are divided into training, validation, and testing subsets with 38623862, 555555, and 937937 videos, respectively. Each video has about 300300 frames on average. The dataset provides ground truth object locations, labels, and object identifications for each frame. Since the annotations for the testing subset has been reserved for the challenge and the evaluation server has been closed, we test on the validation subset as most of the other works.

We use the classical detection evaluation metric for the VID dataset, i.e. the Average Precision (AP) and mean of AP (mAP) over all classes, following the previous works tested on VID .

2 Ablation Studies

We first conduct detailed ablation experiments to study the effectiveness of different components in our method. For fair comparisons, the static image detector baseline mentioned below is a Faster R-CNN network that uses the same settings as we referred to in Section 3.4 except for treating all frames as static images without considering temporal information.

Cuboid Proposal Recall. We first evaluate the recall of proposals by CPN. To do this, we generate a collection of cuboid proposals for each video segment and compute their recall at different IoU thresholds (0.50.5 to 0.70.7) with ground truth cuboids. Fig. 4 shows the quantitative results on the validation set. Firstly, we can see that keeping as few as 5050 proposals already gives reasonably good performance: more than 96.46%96.46\% of the ground truths are recalled for IoU 0.50.5. Secondly, increasing the number of proposals brings only marginal gains for lower IoU thresholds (e.g. 0.50.5) and gives larger gains for higher IoU thresholds (e.g. 0.70.7). The results show that choosing 300300 proposals already achieves satisfactory recall. Thus we only use 300300 proposals for following experiments.

We also show several qualitative results in Fig. 5. The green boxes are the ground truth cuboids and the rest are the proposals generated by CPN. In most cases, there is at least one proposal that has sufficient overlap with the ground truth cuboids, which shows that the CPN can generate reliable cuboid proposals and deals well with videos having single/multiple, small/large, fast/slow moving objects.

Short Tubelet Detection. We then investigate whether the boxes in the short tubelets for short video segments correspond to the same objects. For a testing video, our method first generates a set of short tubelets. Then if all boxes in the short tubelet localize object accurately and correspond to the same object, the tubelet is classified as true positive, and otherwise it is a false positive. After that we compute the mAP. It is obvious that this is a more strict evaluation criterion than the one used for video object detection. Fig. 6 shows the results. We can see that using the strict evaluation protocol only slightly decreases the performance (e.g. from 70.5%70.5\% to 69.8%69.8\% or from 70.1%70.1\% to 68.7%68.7\%), which justifies that the linking results of short tubelets are reasonably accurate.

The Influence of Short Video Segment Length. We discuss the influence of short video segment length. From Fig. 7, we can see that using video segment lengths of 22, 33, and 55 (with T-NMS) all improves over the static baseline. The largest improvement (1.4%1.4\% mAP) is obtained when the video segment length is 22. When the video segment length increases, the performance decreases. In addition, as shown in Fig. 6, we can see that the short tubelet detection performance for long video segments is worse than short ones. There are several reasons explaining this phenomenon. First, longer segments are more probable to generate oversized proposals which have smaller overlap with the ground truth boxes in each frame. Second, the oversized proposals are probable to overlap with the image regions of other objects, causing more ambiguities for accurate localization and classification. Due to the better object/tubelet detection results, we set the video segment length to 22 in the following if not specified. In addition, the video segment length larger than 11 ensures that we can link objects in the same frame.

NMS vs. T-NMS. We study the influence of NMS/T-NMS for object detection. The NMS is implemented by removing boxes for each frame independently instead of removing short tubelets for video segments in the T-NMS. Fig. 7 and Table I show that T-NMS gives better performance than NMS, which confirms that compared with NMS handling each frame independently, simply considering short range temporal contexts contributes to better detection results.

Short Tubelet Linking. Here, we show the improvement by linking short tubelets. We evaluate the performance on slow, medium, and fast ones which are formed according to their speed as done in . We also evaluate the performance on occluded objects (i.e. parts of objects are occluded), following to select 87,19587,195 frames which have more than half occluded objects. As we can see in Table I, compared with the static baseline, considering both short and long range temporal information boosts the performance. When linking objects over the whole video to consider long range temporal context, there is significant improvements (5.4%5.4\% to static and 4.0%4.0\% to without short tubelet linking). Importantly, the performance gains are mainly from the faster objects (6.2%6.2\% for medium and 8.8%8.8\% for fast). It is natural that faster objects may have more variations, thus detecting them depends more on temporal context. Our method also obtains 5.0%5.0\% performance gains for occluded objects, which confirms that our method works well for occlusions. As short tubelet linking performs much better than others, in the following we only report results by short tubelet linking.

Ours vs. Seq-NMS . We compare our results with results by the linking in neighboring frames method Seq-NMS . As shown in Table I, the Seq-NMS obtains better performance than the static baseline, which also confirms the usefulness of temporal contexts. However, the Seq-NMS performs much worse than our method. In particular, the performance improvement for fast moving objects by Seq-NMS is 2.2%2.2\%, whereas our method obtains 8.8%8.8\% improvement. This is because the same object in neighboring frames has different locations and appearances, which influences the quality of object linking, especially for fast moving objects. Thus it is better to link objects in the same frame.

We also combine our CPN and the Seq-NMS for object linking. More precisely, we detect short tubelets from cuboid proposals, and link short tubelets in neighboring frames similar to the Seq-NMS. Results in Table I show that the CPN+Seq-NMS obtains better performance than the method that combines the static detector and the Seq-NMS. This is because our short tubelet detection method can obtain better short tubelets than the Seq-NMS. The CPN+Seq-NMS performs worse than our method, which further demonstrates the effectiveness of our linking in the same frame strategy.

CPN vs. Union Proposal. Here, we compare our CPN with a union proposal baseline. Unlike our method that generates cuboid proposals by CPN, the union proposal method first generates proposals for each frame separately using the static detector, then links proposals in every two neighboring frames according to the proposal IoU, and finally produces the union of linked proposals as cuboid proposals. As shown in Table I (f), the detection performance by the union proposal method is worse than our method. This is because the union proposal method generate cuboid proposals by linking boxes in neighboring frames similar to the Seq-NMS , which cannot obtain high quality cuboid proposals as ours. The cuboid proposal recalls also demonstrate this: 95.1% for IoU threshold 0.5 and 84.9% for IoU threshold 0.7 (union proposal) vs. 96.5% for IoU threshold 0.5 and 90.6% for IoU threshold 0.7 (CPN).

Comparison with the State-of-the-Art Action Detection Method . Finally, we compare our result with the result from which is the state-of-the-art solution in video action detection and adopts the object/action linking framework similar to our method and , by deploying the method on VID. The result by is 60.2%60.2\% mAP which is much weaker than our 74.5%74.5\%. The key point, making our approach perform better, is that the object detection schemes are different. More specifically, our method detects objects for each frame separately, only using the information for the individual frame (with the same proposal for two neighboring frames). localizes action boxes for different frames jointly, thus resulting in poor localization quality. More precisely, stacks features from neighboring frames and uses the stacked features to predict the boxes of these neighboring frames jointly, losing the explicit frame-wise information for predicting the corresponding action box. This is also observed in .

3 Results

We compare our object detection results with the current state of the arts in Table II. First, when only training on the VID dataset, our method obtains the superior result 74.5%74.5\% mAP. To pursue the state-of-the-art detection performance, we follow the previous methods to use the mixture of ImageNet VID and DET datasets for training the detection network, and utilize the standard multi-scale training and testing . As we can see, comparing our 80.6%80.6\% with other methods using the same ResNet-101 network , our method obtains better performance, which confirms the effectiveness of our linking strategy. Importantly, compared with that link objects in neighboring frames, our linking objects in the same frame strategy obtains better performance, which demonstrates that our method can obtain higher quality object linking results.

In particular, the methods in combine feature propagation and the score propagation method Seq-NMS to obtain their results. Feichtenhofer et al. use more anchor scales to obtain better proposals and add a tracking loss to learn better features for performance improvement. There are potential benefits from learning better features in the proposal and detection stages by incorporating other methods such as feature propagation and extra losses into our method.

4 Qualitative Results

Fig. 8 visualizes several detection result comparisons between the static image detector and our method. From the first two rows, we can see that the static method fails to detect the red-panda when there are severe motion blurs and occlusions. This is reasonable because the appearance features have been severely degraded in this situation. After applying the object linking and rescoring, our method successfully classifies the target in the challenging frames. In addition, it is common that the static detectors may confuse with similar classes (e.g., bikes vs. motor-bikes, cats vs. dogs) especially when a frame has low image quality. This problem can also be alleviated by rescoring the detections in the whole video because some frames have correct classifications and can propagate these scores to the challenging frames by object linking.

To show that our cuboid proposal network (CPN) enabling linking in the same frame leads to better localization accuracy, Fig. 9 visualizes several object linking result comparisons between our method and the baseline approach of static detector + linking in neighboring frames over two examples. For (a), we generate per-frame detection results using static image detector and then link detection boxes in neighboring frames by Seq-NMS . For (b), the results are from our approach without later short tubelet linking. One object in each frame in the second to fourth columns has two detected boxes. For (c), the final results are from our approach with short tubelet linking. From the two examples, we can see that the localization accuracy in (a) is poor because of linking in neighboring frames. Here are the analyses. In comparison to the per-frame region proposal network in static detector where the proposals across different frames are independent, the major benefits from CPN include: (1) The proposals are associated. A cuboid proposal consists of two per-frame proposals that are thought to be about the same object, and the resulting detected boxes (predicted for each frame separately) are also thought to be about the same object; and (2) Two nearby cuboid proposals (as well as the resulting detected boxes), e.g., one corresponds to the (n−1)(n-1)-th and nn-th frames, and the other corresponds to the nn-th and (n+1)(n+1)-th frames, are spatially-overlapped in the nn-th frame. Consequently, our approach is able to link the detected boxes in the same frame. The advantage in our linking in the same frame scheme is that we do not need to care about the object movement. In contrast, linking the detected boxes obtained from static detector in neighboring frames might suffer from the object movement and harms the localization accuracy.

5 Runtime

For the case of two frame segments, our method takes 0.350.35s per-frame for testing which is comparable to 0.300.30s by the static baseline. The small extra cost comes from the cuboid proposal generation procedure: a small sub-network processing the two frames separately. The extra time cost is small for the detection stage due to the shared convolutional feature map, the computation time of T-NMS is almost the same as the NMS, and the short tubelet linking is very efficient (about 1010ms per-frame). The speed becomes even faster than the baseline when the short video segment length is larger than 22 (e.g. 0.270.27s and 0.230.23s for video segment length 33 and 55 respectively), because the CPN generates cuboid proposals for all the frames in the segment by computing the features once.

Conclusion

In this paper, we explore to link objects in the same frame for high quality object linking to improve the classification quality. Our method has three main components to achieve our goal: (1) cuboid proposal network, (2) short tubelet detection, and (3) short tubelet linking. Our method obtains the state-of-the-art video detection performance on the VID dataset.

In the future, we will extend our method to handle the two main issues that our approach has: the multi-object issue and the boundary issue. The potential way for the first issue is to generate multiple detection boxes from one proposal. The potential way for the second issue is to recheck the boxes in boundary frames separately. In addition, considering that feature propagation and score propagation are complementary to each other as pointed out in , we will explore how to incorporate feature propagation and score propagation for further performance improvement.

Acknowledgements

This work was supported by National Natural Science Foundation of China (No. 61733007, No. 61572207, No. 61876212), Hubei Scientific and Technical Innovation Key Project, the Program for HUST Academic Frontier Youth Team, and CCF-Tencent Open Research Fund.

References