TF-Blender: Temporal Feature Blender for Video Object Detection

Yiming Cui, Liqi Yan, Zhiwen Cao, Dongfang Liu

Introduction

With the progress of learning-based computer vision, recent research efforts have been extended from image tasks to the more challenging video domains. Video tasks, such as object detection , video instance segmentation , and multi-object tracking and segmentation , hold valuable potentials for real-world applications (i.e., autonomous driving or video surveillance ).

A primary challenge of video object detection is to tackle the feature degradation on video frames caused by camera jitter or fast motion. Under the circumstance, detection algorithms for still images are ill-posed for video tasks. Nonetheless, the video has rich temporal information, on which the same object may appear in multiple frames for a certain time span. The value of such temporal information is explored in prior studies using the post-processing paradigm . These methods firstly perform still-image detection on single frames and then assemble the detection results across temporal dimensions using a disjoint post-processing step (i.e., motion estimation and object tracking). None of the above methods, therefore, operate in an end-to-end fashion. Moreover, if detection on single frames produces weak predictions, the assembling approach cannot improve the detection results. Alternatively, there have been several attempts to boost the performance of video detection using feature aggregation. leverage optical flow to model the feature movement across frames and propagate temporal features to increase the feature representation for detection. With stronger features, the detection results are significantly improved. However, such temporal features are exploited by an intuitive lumping operation, which is oversimplified. In terms of how to organize features in aggregation, we recognize two important predecessors, FGFA and SELSA . Compared to the lumping solution , both methods use similarity scores to select more helpful features for aggregation. The aggregated feature is organized by an adaptive weight at every spatial location for their representations (as shown in Figure 1(a)). Albeit being superior over the prior efforts, FGFA and SELSA encounter several obstacles from achieving optimal performance: 1) They focus on modeling the global relation for every neighboring frame while ignoring the preservation of the local spatial information for aggregation; 2) They primarily consider the global feature relations to the current frames, while having no constraint in feature learning among the neighboring frames (see Figure 1(b)); 3) They take a fixed number of neighboring frames for the feature aggregation, which is heuristic than general.

In this work, we attempt to take a deeper look at video object detection and improve the performance guarantees by organizing temporal information in a more rigorous principle. Inspired by , we propose TF-Blender to organically model features consistently and correspondingly in two ranges. Specifically, we reinforce local similarity in feature space on sequential video frames to depict the continuous and coherence of visual patterns, while identifying semantic correspondence across frames, which makes the temporal representations robust to appearance variations, shape deformations, and local occlusions. In this design, TF-Blender is able to generalize feature aggregation by encouraging the video representation and capturing helpful visual content to improve detection performance. Concretely, we are able to achieve the following contributions:

We propose a framework called TF-Blender, which depicts the temporal feature relations and blends valuable neighboring features to increase the temporal-spatial feature representation across frames.

In TF-Blender, we devise a temporal relation module to manage temporal information and a feature adjustment module to add constraints in feature learning to preserve spatial information during feature aggregation. We, therefore, organize the feature learning between every pair of frames and aggregate features in the whole neighborhood (see Figure 1(c))

Our method is general and flexible, which can be crafted on any detection network. With our novel feature enhancement strategy, we can obtain an absolute gain of more than 0.7%0.7\% in mAP on the ImageNet VID benchmark and 1.5%1.5\% in mAP on YouTube-VIS benchmark for recent state-of-the-arts methods.

Related Works

Video object detection. Different from image object detection, video object detection faces challenging cases (i.e., motion blur, occlusion, and defocus) which rarely occur in images . To handle the challenges in video domains, several works use post-processing techniques on top of still image detectors. For instance, Seq-NMS links bounding boxes across frames with IoU threshold and re-rank the linked bounding boxes; TCN introduces tubelet modules and applies a temporal convolutional network to embed temporal information to improve the detection across frames; T-CNN applies image object detectors to generate results and then uses optical flow to associate the detected results. Although achieving improvements, none of them are trained end-to-end and their performances are still sub-optimal.

Another focus of the recent works is to aggregate temporal features to improve the feature representation for detection. These methods can be divided into three categories: local aggregation, global aggregation, and combination aggregation. Local aggregation methods usually focus on propagating features in a short range on video sequences. Among them, FGFA and MANet are representatives which use optical flow to calibrate and aggregate features across local frames. On the contrary, global aggregation methods rely on long-range semantic information. One seminal work is from SELSA , who computes the semantic similarity between the current frame and its neighbours across the whole video in order to perform temporal feature aggregation. Different from the methods which exploit features locally or globally, MEGA introduces a memory module to use both local and global features to enhance the visual representation of the current frame. The aggregation methods achieve further performance gain over the post-processing methods, but they generally focus on higher-level video frame selection instead of exploring lower-level temporal features exploitation.

Video instance segmentation. Similar to video object detection, MaskTrack R-CNN extends instance segmentation from image domain to video domains which requires segmenting and tracking instances across frames. However, most of the current methods like MaskProp , EnsembleVIS focus on how to track instances across frames rather than how to generate high-quality features for detection, segmentation, and tracking. In this work, we, therefore, propose a more principled solution, which effectively transforms and exploits valuable temporal features for the video object detection task.

2 Relation Learning

Relation learning is widely used for different tasks (i.e., point cloud analysis and image understanding ) to describe the relationship between the current feature and its neighbors. RS-CNN extends regular grid CNN to capture local point cloud features using geometric topology constraints among points. Similarly, PointConv models the feature relation by computing both the local coordinates and point cloud density. Both methods capture local features in geometric space. On the contrary, DGCNN defines EdgeConv which captures local point relation in high-dimensional feature space and updates the neighborhood for the kernel dynamically at each layer. Similarly, some recent works attempt to leverage relation learning for object detection. Inspired by which proposes an object relation module for still image object detection, RDN introduces a relation distillation network to aggregate features based on object relation to improving the features for video object detection. MEGA extends the relation learning from RDN and proposes a memory-enhanced global-local aggregation network, which organically manages long-range (global) features and short-range (local) features for aggregation in order to increase the feature representation of current time for detection. However, the focuses of the above methods are the selection of higher-level video frames for aggregation rather than modeling lower-level temporal relation to increasing the feature representation. Different from these methods, we propose a more general approach for relation learning in feature aggregation. Our TF-Blender can robustly depict the salient correspondences between the feature of the current frame and neighboring frames and exploit only valuable features for a stronger detection.

TF-Blender

The conventional feature aggregation methods generally work in a constrained fashion. Given a set of neighboring frames Fj\textbf{F}_{j} of the current frame Fi,∀Fj∈N(Fi)\textbf{F}_{i},\forall\textbf{F}_{j}\in\mathcal{N}\left(\textbf{F}_{i}\right), their corresponding features fjf_{j} are weighted equally based on the feature similarity to Fi\textbf{F}_{i} in order to aggregate the temporal feature Δfi\Delta{f}_{i}:

The principal problem of feature aggregation, therefore, is to calculate weights wijw_{ij} and select representative neighboring feature fjf_{j}. Different from the above simple paradigm, we exploit the temporal features from a general perspective. To achieve this goal, our TF-Blender crafts on three novel architectural modules, temporal relation module, feature adjustment module, and feature blender module, to boot the detection performances (see Figure 2).

2 Temporal Relation

where gg is a feature relation function to describe the temporal relation between fif_{i} and fjf_{j} and M\mathcal{M} is a masking function to calculate the adaptive weight based on gg. As shown in Figure 3(c), our temporal relation can enhance the feature representations from the region of interest and suppress the irrelevant features. More concretely, we compute M\mathcal{M} in Eq. 2 using a mini-network (see Figure 4). Compared with the CoefNet in LMP , our feature adjustment module is built on a lighter architecture, which makes our TF-Blender computationally efficient. The input of the module is fif_{i} and fjf_{j}, marked as red and blue cuboids respectively. Feature relation function gg describes the relation between fif_{i} and fjf_{j} and generates the input (the gray cuboid) of the mini-network M\mathcal{M}. Afterward, we apply three convolution layers (the yellow cubes) to generate the final adaptive weights W(fi,fj)\mathcal{W}\left(f_{i},f_{j}\right) (the purple cuboid). The selection of the feature relation function gg will be discussed in 4.1.

3 Feature Adjustment

Our feature adjustment module aims to represent the feature consistency and salience of the neighboring frames for feature aggregation. A simple solution is to directly use feature fjf_{j} from frame Fj\textbf{F}_{j} as the follow:

However, fjf_{j} cannot be guaranteed to be valuable for aggregation as there is no constraints between these neighboring features. Therefore, we aggregate every neighboring frame feature fjf_{j} before aggregating the current frame feature fif_{i}. We get feature representative F(fi,fj)\mathcal{F}\left(f_{i},f_{j}\right) by aggregating fjf_{j} with the other neighboring features fm,∀Fm∈N(Fi),Fm≠Fjf_{m},\forall\textbf{F}_{m}\in\mathcal{N}\left(\textbf{F}_{i}\right),\textbf{F}_{m}\neq\textbf{F}_{j} (see Figure 2). During feature adjustment, we use the temporal relation module to generate adaptive weights for neighbouring feature aggregation and the process can be expressed as:

where ⊗\otimes is element-wise multiplication, fmf_{m} is the feature of the neighboring frame except itself, and W(fj,fm)\mathcal{W}\left(f_{j},f_{m}\right) is Eq 2, which can be expressed here as:

4 Feature Blender

In our feature blender module, we first enhance the results of the temporal relation module with the non-linear function ReLU so that the contrast between the area of interests and background can be captured (see the blender module in Figure 2). We formulate this process as:

Meanwhile, we normalize the results of the feature adjustment module with the softmax function over all the channels to improve the generalization of our model. On the top of the feature blender module in Figure 2, blue dots are normalized to green dots by the softmax function with the guidance of purple double arrows. The process can be expressed as:

In our feature blender module, we force W^(fi,fj)\hat{\mathcal{W}}\left(f_{i},f_{j}\right) to be 00 if the adjusted neighboring feature is very similar to the feature of the current frame, shown as dashed purple double arrows in the feature blender module part of Figure 2. We use the cosine distance to describe the similarity between F^(fi,fj)\hat{\mathcal{F}}\left(f_{i},f_{j}\right) and fif_{i}. If the cosine distance is bigger than δ\delta, W^(fi,fj)\hat{\mathcal{W}}\left(f_{i},f_{j}\right) is force to be 00. We define this process as:

We have this design because most of the current feature aggregation-based methods have a fixed number of neighboring frames in aggregation. However, for neighboring frames which include issues of severe motion blur or defocus, aggregating them are irrelevant and redundant, which may cause unwanted ambiguity.

Finally, we use element-wise multiplication to combine the results of from Eq. 7 and Eq. 8 to perform the feature aggregation:

Experiments

Evaluation metrics. Following , we report all results on using the mean average precision (mAP). Video object detection setup. We evaluate our methods with MEGA , SELSA ,FGFA , and RDN , the three state-of-the-art systems. We perform our training and evaluation on the ImageNet VID benchmark , which contains 3,862 videos for training and 555 videos for validation. Following the widely used protocols in , we train our model on a combination of ImageNet VID and DET datasets. We implement our method mainly based on the source code of the original method. The whole network is trained on 8 RTX 2080Ti GPUs with SGD. During the training and inference process, each GPU holds on one set of images or frames. During the training process, the encoder parameters are frozen and an NMS of 0.5 IoU is adopted to suppress detection redundancy.

Video instance segmentation setup. We also evaluate our proposed method with state-of-the-art MaskTrack R-CNN and SipMask . We perform our training and evaluation on the YouTube-VIS benchmark , where there are 3,471 videos for training and 507 videos for validation. During the training process, we use weights pretrained on MS-COCO and use 8 RTX 6000 GPUs with SGD. In both training and evaluation, the original frame sizes are resized to 640×360640\times 360.

Parameters. For mini-network M\mathcal{M} in Eq. (2), a three-layer CNNs is introduced to adapt the channels for feature aggregation. Feature relation function gg is defined as a concatenated tensor of fif_{i}, fjf_{j}, fi−fj,fj−fif_{i}-f_{j},f_{j}-f_{i} and the δ\delta in Eq. (8) is set to 0.7.

2 Main Results

Results on ImageNet VID benchmarks. We compare state-of-the-art systems crafted on our method with their original implementations. For a fair comparison, we used the codes provided by the original papers and re-implement them with our proposed method. The results are demonstrated in Table 1. Based on the results, our proposed methods substantially improve the performance of every compared method listed in the table with the same backbone. For head-to-head comparisons, all the methods with the same backbone can leverage our proposed methods to improve their performances on detection results around 0.7%0.7\%-1.5%1.5\% on accuracy. Among them, FGFA with our proposed method has the highest improvement compared with other methods. Among them, local aggregation and global aggregation methods like FGFA and SELSA can have a better improvement with our proposed methods compared with combination aggregation methods like RDN and MEGA . We argue that the limited performance gains come from the combination aggregation methods, which consider both local and global features and make detection more robust to issues like motion blur in videos. Figure 5 shows some examples of detection results with our methods integrated. Based on the examples, we can see that our proposed method can help solve the problem of weak detection with rare pose and part occlusion situations.

Experiments on YouTube-VIS benchmark. We also evaluate our proposed method on YouTube-VIS dataset and report our results on the validation as . Most of the current video instance segmentation methods focus on how to generate high-quality masks and link the same objects across frames with features extracted by the backbones like ResNet while only a few of them pay attention to improve the features for mask generation and object tracking. We add our proposed methods to these video instance segmentation methods to evaluate the effectiveness of our TF-Blender on issues like motion blur and defocus in videos. The results with ResNet-50 as backbones are shown in Table 2. From Table 2, our proposed methods achieve competitive results under all evaluation metrics. With our proposed methods, MaskTrack R-CNN and SipMask can be improved by more than 1.6%1.6\% on the AP metric. The bottom part of Figure 5 shows an example of detection and segmentation results with our integrated.

3 Ablation Study

We carry out extensive ablation studies to discover the optimal settings related to different settings of our system using FGFA .

Analysis of contributing components. We first conduct experiments on the effect of every component in our proposed method and the results are shown in Table 3. The baseline model a is the original FGFA. Every component of our proposed method (temporal relation, feature adjustment, and feature blender) contributes towards improving the overall performance in detection accuracy. By introducing the temporal relation module, the performance of model b can be improved by 0.7%0.7\%. Model c adds our feature adjustment module to the baseline and gets an improvement of 0.3%0.3\% compared with the baseline model a. We add our feature blender module to model a to generate dynamic numbers of neighboring frames for feature aggregation and get model d, which is 0.5%0.5\% better than the original model on mAP metric. Model e, f, and g come from the combination of models a, b and c. As can be shown in Table 3, by combining every two of our proposed methods, the video object detection performance can be further improved. Compared with the baseline model a, our full model h can obtain an absolute gain of 1.5%1.5\% in accuracy of video object detection.

Analysis of temporal relation. We conduct ablation studies on the choice of gg in Eq. (2). During these experiments, all the other experimental settings are kept the same. We first try different combinations of fif_{i} and fjf_{j} for gg on FGFA as Table 4. A naive idea is to use just fif_{i} and fjf_{j} as input and there is 0.5%0.5\% improvement on FGFA. We think that the performance is limited because only individual frame features are taken into account which is not enough to describe the relationship between the fif_{i} and fjf_{j}. Thus, we introduce the difference between fif_{i} and fjf_{j} to gg and get an improvement of 0.8%0.8\% for FGFA. We then use the summation of fif_{i} and fjf_{j} as gg to generate W(fi,fj)\mathcal{W}\left(f_{i},f_{j}\right) but there is only 0.1%0.1\% improvement. We also make a combination between fi+fjf_{i}+f_{j} with the other choices mentioned above (like fi,fjf_{i},f_{j}, and fi−fjf_{i}-f_{j}), but the results of the combination are worse than those of the original. We think the reason why fi+fjf_{i}+f_{j} is not suitable to describe the relations between fif_{i} and fjf_{j} is fi+fjf_{i}+f_{j} works like an average filter which mixes the pixels with higher responses and those with lower responses in the feature map. Besides the experiments mentioned above, we also try fi,fj,fi−fjf_{i},f_{j},f_{i}-f_{j} and get an improvement of 1.1%1.1\%. Finally, we choose fi,fj,fi−fj,fj−fif_{i},f_{j},f_{i}-f_{j},f_{j}-f_{i} as our feature relation function gg, which has the highest detection accuracy. Since fif_{i} and fjf_{j} denote the current and adjacent features respectively. Frame Fj\textbf{F}_{j} could be a frame before or after the current frame Fi\textbf{F}_{i}. Thus, it is imperative to calculate both fi−fjf_{i}-f_{j} and fj−fif_{j}-f_{i}, as they model the different temporal correspondence and consistency.

Experiments on M\mathcal{M}. We conduct experiments on the design of M\mathcal{M} for the temporal relation module, especially on the number of layers of M\mathcal{M} for the mini-network. Model a is the simplest design where there is only one convolution layer with kernel size 1×11\times 1. By keeping the kernel size fixed and adding one more convolution layer, model b can increase the mAP by 0.2%0.2\%. When there are three convolution layers with kernel size 1×11\times 1, the detection accuracy can obtain 79.2%79.2\% as model c. However, when adding more convolution layers, as in model d, the detection accuracy begins to decrease. We argue that the increasing number of convolution layers introduces arduous parameters in the mini-network which cause overfitting. In model e, we change the kernel size from 1×11\times 1 to 3×33\times 3 and get an improvement of detection accuracy by 0.1%0.1\%.

Analysis of object sizes and motion speeds. We also investigate the effect of our TF-Blender on the object sizes and motion speeds of the objects. We use the same definition as MS-COCO and FGFA for object sizes and motion speeds respectively. We use mAP as evaluation metrics and visualize the improvement of performance on objects with different sizes and motion speeds as Figure 6. We notice that our method has different improvements on objects with various motion speeds. As shown in Figure 6 (a), there is a higher improvement for objects with slow motion speeds compared with those with fast and medium speeds. We think that there may be two reasons. One is that even though our proposed method can help improve the detection accuracy for objects with fast motion speeds, it’s still a challenge to have accurate enough detection results for all the objects with fast-motion speed. Another reason is that objects with slow-motion account for 37.9%37.9\% in ImageNet VID benchmark while those with medium and fast motion speeds are 35.9%35.9\% and 26.2%26.2\% respectively.

Another critical observation from our experiment that our method can offer the highest improvement for detection on large objects, as shown in Figure 6(b). This resonates with the assumption of our proposed method: since large objects have larger feature map sizes, the corresponding pixel can benefit more from an individual weight for fine-grained feature encoding. For small objects, since their feature maps are small, the weights for aggregation have less contribution to feature representation improvement.

Speed-accuracy tradeoff. The computational loads for convectional methods (i.e., FGFA and SELSA ) stem from two major sources: 1. feature extraction (encoding) network Nex\mathcal{N}_{ex}; 2. task network Ntk\mathcal{N}_{tk}. Thus, the runtime complexity for the above methods is:

While the proposed TF-Blender approach is adopted, the computational cost can be defined as:

where Ntf\mathcal{N}_{tf} is the cost for the TF-Blender module and ii is the number of aggregated frames. Typically, O(Ntk)≪O(Nex)\mathcal{O}\big(\mathcal{N}_{tk}\big)\ll\mathcal{O}\big(\mathcal{N}_{ex}\big) and O(Ntf)≪O(Nex)\mathcal{O}\big(\mathcal{N}_{tf}\big)\ll\mathcal{O}\big(\mathcal{N}_{ex}\big). Thus, the cost ratio rr can be expressed as:

This increasing computational cost is affordable because the impact of i⋅O(Ntf)i\cdot\mathcal{O}\big(\mathcal{N}_{tf}\big) is negligible.

We visualize the speed-accuracy tradeoff of FGFA as an example (cf. Figure 7). With the increasing number of input frames, FGFA with TF-Blender achieves significant improvement in accuracy while the runtime increase keeps in an affordable range.

Conclusion

In this paper, we discuss the problems of video object detection and introduce a framework named TF-Blender which contains temporal relation, feature adjustment, and feature blender modules to solve the problem of feature degrading in the video frames. Our method is flexible and general, which can be adopted by any learning-based detection network to achieve improved performance. Extensive experiments demonstrate that, with the integration of our proposed method, the current state-of-the-art methods can improve video object detection accuracy on ImageNet VID and YouTube-VIS benchmarks by a large margin. We believe that our TF-Blender can be a valuable addition to the existing methods for temporal feature aggregation for video detection and TF-Blender can be extended to other video analysis tasks like video instance segmentation.

References