InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges
Guo Chen, Sen Xing, Zhe Chen, Yi Wang, Kunchang Li, Yizhuo Li, Yi Liu, Jiahao Wang, Yin-Dong Zheng, Bingkun Huang, Zhiyu Zhao, Junting Pan, Yifei Huang, Zun Wang, Jiashuo Yu, Yinan He, Hongjie Zhang, Tong Lu, Yali Wang, Limin Wang, Yu Qiao
Introduction
Ego4D is the latest large-scale egocentric video understanding dataset presented by Facebook AI Research (FAIR). Different from the previous video understanding datasets, Ego4D unintentionally collects persistent 3D egocentric data, including the camera wearer’s physical surroundings, interactive objects and actions, and high-level social behaviors. It aims to catalyze the next era of research in first-person visual perception. The accompanying 5 benchmark tasks open new research directions and stimulate research broadly from around the world to a large extent. These benchmark tasks cover the basic components of egocentric perception—indexing past experiences, analyzing current interactions, and predicting future activities. To promote Ego4D and explore its research methodologies, several challenges are organized.
In the Ego4D ECCV2022 Challenge, we joined in five tracks Moment Queries, Natural Language Queries, Future Hand Prediction, State Change Object Detection, and Short-term Object Interaction Anticipation. As shown in Table 1, we won all championships in these tasks. We share our solutions in this technical report.
Despite efforts to adapt cutting-edge task heads to our participated tracks, we leverage our developed InternVideo, a video foundation model, to support our competitions. As it is an effective paradigm to exploit a strong foundation model to address downstream tasks, simplifying head designs. InternVideo involves both masked autoencoder and multimodal learning, and we chiefly employ its two components VideoMAE and UniFormer . VideoMAE offers a spatio-temporal representation using a vision transformer encoder by a masked video reconstruction pretask, while UniFormer integrates local spatio-temporal modeling into transformers for efficient video representation learning. Besides of these two backbones, we also explore others, e.g. Swin, EgoVLP, ResNet, and more, for comparisons or fusion.
In the remaining of this report, we will introduce the related work first. And then we will detail our solutions along with experiments to each joined Ego4D track, after briefing the employed VideoMAE and UniFormer. Finally, we discuss the limitations of our work and conclude this paper.
Related Work
3D Convolutional Neural Networks (CNNs) have been dominant in video understanding . Due to the difficult optimization problem and large computation of 3D convolution, great efforts have been made to factorize 3D convolution. R(2+1) , P3D and TANet divide the 3D convolution into 1D temporal convolution and 2D spatial convolution, while CSN , X3D and MoViNet propose 3D channel-separated convolution. However, 3D convolution struggles to capture long-range dependencies because of the local receptive field . Inspire by the success of Vision Transformers (ViTs) in the image domain , researchers try to apply global attention for spatiotemporal modeling . For efficient and effective video understanding, MViT introduces a hierarchical structure with pooling self-attention, VideoSwin extends window self-attention to 3D space and UniFormer proposes to unify convolution and self-attention for better accuracy-computation trade-off.
2 Temporal Action Localization
Temporal Action Localization (TAL) is aimed at detecting the boundaries and categories of the action segments in untrimmed videos. Current temporal action detection methods can be divided into one-stage and two-stage methods. The two-stage methods decouple the tasks of generating proposals and classifying actions. For instance, use a flexible way called boundary matching to generate high-quality proposals. They predict each frame’s start and end confidence, then match the frames with high start and end confidence to generate the proposals and evaluate their confidence. The one-stage methods generate action proposals and corresponding action labels simultaneously in a single model. Recently, explores the end-to-end training methods and yields outstanding performance. However, considering the end-to-end training overhead for video data, utilizing a feature-based one-stage detection method is a more efficient and convenient option. Among these methods, VSGN adopts GNN to aggregate multi-scale temporal features to generate action predictions based on dense anchors. ActionFormer builds a transformer network to predict the offsets of the start and end position through a well-implemented anchor-free mechanism. In the Ego4D challenges, the Moment Queries (MQ) task aims to query the specific moments consistent with the TAL task. We adopt VSGN and ActionFormer as our detection heads to validate the performance of VideoIntern on this task.
3 Video Temporal Grounding
Video Temporal Grounding (VTG), which needs to retrieve video segments using natural language queries, was introduced in . Early works explored how to utilize text queries. One method is metric learning based, such as , which uses metric learning loss functions with the distance as the similarity measurement to match the given text queries to the right video moment. The other is detection-based, transforming the text query as a dynamic filter as a condition for extracting the temporal feature. It can inherit excellent mechanisms from temporal action detection and object detection. VSLNet adopts the context-query attention module to predict video segments corresponding to the text queries. In the Ego4D challenges, the Natural Language Queries (NLQ) task aims to localize the correct temporal segments through natural language queries, consistent with the VTG task. We use VSLNet as our grounding head to validate the performance of VideoIntern on this task.
4 Spatio-Temporal Action Localization
The purpose of Spatio-Temporal Action Localization (STAL) is to predict people’s location in keyframes and classify the person’s ongoing actions in videos. At present, most methods divide this task into two subtasks. They first generate human bounding boxes in keyframes and then extract 3D ROI features for these boxes to classify actions. The generated people boxes by Faster R-CNN is widely used in current STAL methods. Considering that there may be various actors or objects in a keyframe, these methods, such as , focus on capturing the relations and contexts in different levels between the objects in the video.
ViT uses the global self-attention mechanism to mix the spatial patches. replaced spatial self-attention blocks with Spatio-temporal self-attention blocks so that ViT can achieve cube-level dense spatiotemporal interaction. Due to the full spatiotemporal interaction implemented in the backbone, the well-designed interaction modules in the STAL methods become unnecessary. Compared with the STAL task, the Short-term Object Interaction Anticipation (STA) task in Ego4D challenges has a similar form but different predicting requirements. Given a pre-condition clip, it needs first to forecast the bounding boxes of the objects that will be interacted with, then forecast the noun category of objects, the verb category of the interaction, and the time to contact objects. We use a similar manner to Spatio-temporal action localization to complete this track.
5 Object Detection
Object Detection is a classic 2D computer vision task to predict the bounding box regression and pixel-level classification results. The traditional CNN-based object detectors have been widely studied in the past years, such as R-CNN , Fast R-CNN , Faster R-CNN , and so on. In recent years, transformer networks have become popular, and detectors based on transformers have emerged as the times require. DETR and deformable DETR used the transformer decoder to perform end-to-end object detection. AdaMixer proposed a fast-converging decoding module in query-based detector based on sparse sampling and dynamic MLP mixer. DINO made a series of optimizations and achieved state-of-the-art performance on the COCO dataset.
Methodology
We choose CSN, VideoMAE , and UniFormer as our backbones for feature extraction. These three backbones have different architecture designs: pure convolution network, pure transformer network, and convolution-transformer hybrid network. We hypothesize that the representations from these three backbones are different and complementary. We employ different task-specific heads to complement different tasks.
The used backbones have been pre-trained on action recognition datasets already. Specifically, CSN , VideoMAE , and UniFormer are pre-trained on IG65M , K700 , and K600 , respectively.
These backbones need further finetuning as we find there is a distribution gap between data used in the pre-trained datasets and egocentric ones. Using the pre-trained backbones directly would lead to performance degradation in the egocentric video. To alleviate the negative effects brought by this gap, we finetuned these backbones on the Ego4D training set. Specifically, we adopt the clip-level annotations of EgoVLP on the training set for rapid development.
In training on the Ego4D dataset, we finetune two model variants for each backbone using two types of annotations, considering that Ego4D has both the verb and noun category annotations for short clips in the video. Intuitively, models trained by the verb and noun annotations can capture action and scene information, respectively. Taking the verb annotation as an example, we preserve all the videos that contain the verb annotations. For videos containing multiple verb categories, we treat them as single-category annotations for multiple videos (the same video). To construct rapidly a validation set to evaluate the performance of fine-tuning, we sample 5% of the videos of each class as the validation set. Finally, we use the standard single-class action recognition training method to finetune the backbones with our constructed verb and noun datasets.
2 Experiments
We use ir-CSN-152, VideoMAE-L, and UniFormer-B as our backbones. These models are trained for 10 epochs by AdamW optimizer and cosine schedule. During fine-tuning, we set the batch size to 256 and the learning rate to 5e-4. The other fine-tuning setting is shown in Table 1.
We train these models on the verb and noun subsets and use the Top-1 and Top-5 accuracy of the validation set to measure the quality of finetuning. The results are shown in Table 2. Due to its high performance, we mainly adopt VideoMAE-L as the feature extractor for the two temporal localization tasks, NLQ and MQ.
Track 1: Moment Queries
The Moment Queries track serves as the Episodic Memory task to query the moments of some high-level activities or “moment” names (that can be transformed into discrete labels). For an untrimmed egocentric video, we denote it as , where indicates the length of the video and is the -th frame. We denote the temporal annotation of action instances as in the video which has instance. , , and are the start, end boundary, and categories of the instance , respectively. The model generates predicted moment segments that should cover with high recall and high temporal overlapping.
2 Approach
Due to the length of each untrimmed video coverage ratio of the activities being large in the MQ dataset, we adopt the two-stage method to localize temporal segments. Specifically, we first finetune our backbone with MQ labels and then extract the pooled Spatio-temporal features as the input of localization heads. For validating comprehensively the feature extracted by our backbones, we use multiple localization heads as our candidate task-specific heads to observe the experimental results.
Fine-tuning Settings. Since the gap between different datasets, there are multiple routes for transfer learning. We attempt the following two methods that fine-tune backbones with MQ labels. One-stage: The first one is one-stage finetuning. We skip the full Ego4D dataset and directly finetune ir-CSN-152 pre-trained on IG65M or VideoMAE-L pre-trained on K700. We denote it as “K700 MQ”. Two-stage: The second method finetunes the backbones on the full Ego4D dataset and then continues finetuning the backbone on the MQ sub-dataset.
Feature Extraction. We train the temporal localization methods with offline video features instead of end-to-end training. We extract Spatio-temporal features on videos through a sliding snippet approach at 15 FPS. CSN is adopted to extract a 2048-dimensional feature vector for each snippet. Each snippet contains consecutive frames with snippet interval . For VideoMAE, we extract a 1024-dimensional feature vector for each snippet that contains consecutive frames with interval . Further, we separately extract verb and noun features for all backbones and aggregate them to enhance video representations.
Localization Head. We adopt the official baseline VSGN for this temporal action localization task with MQ. Then we replace its detector with ActionFormer , further improving the final localization performance.
3 Experiments
We perform experiments with different backbones, fine-tuning settings, and localization heads on the validation set and submit inference results on the testing set to EvalAI’s test server.
We first evaluate the localization performance with the one-stage fine-tuning setting and VSGN. We conducted some experiments with this method, and the results are listed in Table 3. Comparing VideoMAE features and CSN features, we find that the former can bring higher mAP, and the latter can achieve higher Recall on the test set. We also explore the spatial-level Multi-View Fusion (identified as “MVF” in the table) method to enhance temporal feature representations, improving localization performance but increasing computational overhead.
Finally, we adopt the two-stage fine-tuning method as our solution for MQ tasks. We extract the two types of features (denoted as “K700 Verb MQ” and “K700 Verb” in the table) with and without the second finetuning stage to explore their effect on localization performance. ActionFormer is used to replace VSGN to unlock the potential of the temporal features. The experiment results on the validation and test set are shown in Table 4.
Comparing Table 4 and Table 3, introducing the full Ego4D dataset greatly boosts mAP and Recall. As shown in the second and third rows in Table 4, the second stage of finetuning transfers the feature representation of the backbone from the lower-level representation of verbs or nouns to the higher-level activity representation of the MQ dataset. It improves the overall localization performance, with +5.55 in Recall and +4.28 in mAP on the test set.
Track 2: Natural Language Queries
The Natural Language Queries track serves as the Episodic Memory task to query the moments corresponding to some text. For an untrimmed egocentric video, we denote it as , where indicates the length of the video and is the -th frame. We denote the temporal annotation of action instances as in the video which has instance. , , and are the start, end boundary, and text query corresponding to the instance , respectively. The model generates predicted moment segments that should cover and high temporal overlapping.
2 Approach
Benefiting from the high performance of the VideoMAE feature in the MQ task, we explore its performance in the NLQ task. We adopt the method mentioned in Sec 4.2 to extract video features. The official baseline VSLNet is applied to solve this task.
3 Experiments
Using EgoVLP as our baseline, we perform several further experiments shown in Table 5. Different configurations are identified by capital letters from “A” to “F”. We will introduce these configurations and the improvements they bring.
B & C. We first replace the video feature of EgoVLP with our verb and noun features, respectively, and preserve the text encoder of EgoVLP. Comparing EgoVLP, our video features improve by about +2.0 R1@0.3 and R1@0.5. Meanwhile, the noun feature is better than the verb feature, which suggests some domain gaps between the verb feature and the noun feature.
D. Based on the previous experiment, we assume that the verb and noun features pay more attention to the motion and scene information in videos, respectively. We fuse the verb and noun features before feeding them into the model. It brings at least a +1.0 R1 improvement, proving that verb and noun features are complementary.
E. We further tried to fuse the video features of EgoVLP and gained about +0.9 R1@0.3 and R1@0.5 improvement on the test set. Since EgoVLP is a multi-modal pre-training method, it may be complementary to full-supervised pre-training.
F. Finally, we post-fuse the predictions of D and E to improve the overall NLQ performance.
Track 3: Future Hand Prediction
The Future Hand Prediction track serves as the forecasting task to forecast the spatial location of future hands. Specifically, we denote the contact frame as , the pre-condition frame as , and the three frames preceding the pre-condition frame by , , and as , , , respectively. Formally, given an input egocentric video before the pre-condition time step (denoted as , with referred as observation time), this task seeks to predict the positions of both hands in the future key frames, where .
2 Approach
We follow the baseline method, predicting 20 categories for short-term (1.5 seconds) historical information. The 20 categories indicate the future’s absolute spatial coordinates of five pairs of hands. We use L1 loss to regress these coordinates.
Inspired by the locality introduced by the convolution module in object detection benefits the regression boundary, we adopt UniFormer as our backbone, which contains depth-wise convolution modules to preserve the local spatial information explicitly. This historical spatial local information is a major source for regressing the coordinates in the future.
3 Experiments
We attempt to use VideoMAE-L and UniFormer-B to regress the spatial location of the future hands. The results are shown in Table 6. We first perform experiments using VideoMAE-L that only contains global self-attention modules to ablate the effect of convolution modules. When training VideoMAE-L, the network is difficult to converge, and the final result is weaker than I3D .
We use UniFormer-B as the backbone of our final solution. The forecasting results have been much better when inputting 4 frames with 320320 resolution than I3D. It may be because of higher spatial resolution or a stronger convolution hybrid backbone. Furthermore, we increase the number of input frames and fuse multi-view predictions to enhance the regression performance.
Track 4: State Change Object Detection
State Change Object Detection (SCOD) is the task of detecting the object undergoing a state change from the given egocentric video clips. Specifically, each given video consists of three temporal frames, i.e. precondition (PRE), point-of-no-return (PNR), and post-condition (POST), and the goal is to predict the 2D bounding boxes of the state change object in the PNR frame.
2 Approach
For this task, the officials provide several single-frame baselines covering a wide range of detection frameworks, including CenterNet , Faster R-CNN , and DETR . These methods adopt ResNet and DLA as the backbones and achieve decent results. Nevertheless, we argue that the above methods still have some room for improvement due to the sub-optimal choice of the backbone network and detection head.
In this technical report, we develop a stronger detector for the SCOD task. Specifically, we follow the official baseline to build a single-frame detector and perform 2D object detection on the PNR frame. Our method consists of two key components: (1) an image backbone (e.g., UniFormer-L or Swin-L ), and (2) a query-based detection head DINO .
On this basis, we further explore the transfer learning from general object detection to egocentric images, including three different pre-training tasks of ImageNet classification , COCO detection , and Objects365 detection (i.e., a larger detection dataset). The experiments indicate that our method achieves a promising improvement over the official baselines, and the pre-training on general object detection can derive significant benefits for the SCOD task. We hope this method can serve as a strong baseline for egocentric object detection.
3 Experiments
Our detection experiments are based on the SCOD dataset and the MMDetection codebase. We adopt UniFormer-L pre-trained on ImageNet-1K, or Swin-L pre-trained on ImageNet-22K as the backbone. In addition, we employ DINO as the detection head, in which the numbers of content and denoising queries are fixed to 900 and 1000, respectively.
Firstly, we tried to train the SCOD dataset without extra detection datasets. Further, we study the transfer learning from two general object detection datasets (i.e., COCO and Objects365 ) to the SCOD dataset. The pre-training schedules for these two datasets are 12 epochs and 26 epochs, respectively. During SCOD fine-tuning, the shorter side of the input image is resized between 800 and 1600, while the longer side is at most 2000. All models are trained with AdamW optimizer (batch size of 16, initial learning rate of 110-4, and weight decay of 0.0001) for 12 epochs.
As shown in Table 7, when using only ImageNet-22K pre-training, our method yields an impressive score of 28.0 AP on the SCOD validation set, outperforming previous official baselines by at least +12.5 AP. We can also see that pre-training on general object detection can greatly benefit the SCOD task. For instance, COCO pre-training promotes the detection performance to 32.2 AP, and Objects365 pre-training further achieves a big jump to 36.4 AP. We submitted the best result to the test server, achieving 37.2 AP on the test set and ranking 1 on the leaderboard.
Track 5: Short-term Object Interaction Anticipation
The Short-term hand object prediction task aims to predict the next human-object interaction happening after a given timestamp. For a given video, we need to detect the spatial location of active objects and perform noun classification. These active objects will be considered to interact with people at the time in the future. The action category of the interaction also needs to be identified, and the value of the needs to be predicted.
2 Approach
In this task, the baseline method uses a two-stage approach, which is commonly used in ST-AL tasks. Given a series of video keyframes , these keyframes can be extracted from the same video or different videos. In the first stage, the object detector is trained first, and the detector generates a series of localization boxes from the keyframes of the video. We then predict a noun category for each bounding box. In the second stage, for each keyframe, we sample a clip before it to get a series of frames and feed them into the backbone to extract features. In the final classification stage, we use RoIAlign to capture areas of interest from the extracted features, classify actions, and regress contact time based on these features.
In the baseline method, Faster R-CNN is used for detection in the first stage. In the second stage, the baseline uses SlowFast to perform feature extraction on the input video and classify it at the end. As we mentioned before, this task needs to predict actions and time to contact, both of which are related to the interaction between people and objects. However, the receptive fields of the convolutional neural network are limited, and capturing people’s interactions with things is difficult. In addition to the limitation in capturing the interaction relationship, the baseline method directly uses RoI features to predict time to contact, which is an indication of the time of human-object interaction and is related to the location where the human and the object are located. It is not reasonable to use RoI features only to predict the time.
Based on these considerations, we have made the following improvements:
(1) We use the ViT with space-time attention pre-trained by VideoMAE to capture the interaction between people and objects. It exploits self-attention with long-range modeling capability to characterize such an interaction, which benefits action recognition.
(2) To better combine the time to contact with the position of the box, we perform a positional encoding operation on the boxes and fuse the positional information of the box with the RoI feature to predict the time to contact.
(3) Although the bounding boxes provided by Ego4D officials already achieved decent performance, we used our own trained detectors in the first stage to further improve the prediction quality and then conducted some post-processing to reduce redundant boxes.
We used the open-sourced detection framework MMDetection to train the current state-of-the-art detector DINO . DINO is an improved object detection algorithm based on DETR. It uses a transformer network to generate a fixed number of bounding boxes per keyframe. In this task, we generate 900 boxes for each keyframe and then perform the non-maximum suppression (NMS) post-processing to remove redundant predictions.
2.2 Video Backbone
We used VideoMAE as the backbone, which is a complete transformer network. VideoMAE is pre-trained on the K700 dataset and then trained for basic classification on the Ego4D dataset. The Transformer-based network is chosen because of its explicit attention mechanism and competitive performance in various discriminative tasks. It is effective for modeling the long-range relationship between people and objects and is conducive to the recognition of human actions. After the pre-training of VideoMAE, we finetune the encoder to extract features from the input video in the whole pipeline.
Although we already have well-performing boxes in the first stage, these boxes are still quite different from the ground truth. In the original configuration of the baseline method, only the boxes predicted in the first stage are used for training. Since the boxes are not accurate enough, it is highly likely to bring inevitable errors to the second stage, degrading its training. The Spatio-Temporal localization task meets a similar problem as well. The common practice is using the ground-truth boxes for training and the boxes predicted by the detector for validation and testing. However, due to the small number of ground truth boxes in each frame in the training set, the amount of training data will be greatly reduced. To improve the generalization ability of the model, we finally adopt the method of taking both ground truth boxes and predicted boxes as input. We find that it produces better results than using either ground truth or predicted boxes alone.
We initially followed the baseline method to directly perform the RoIAlign operation on the input box to obtain the RoI feature, and predict the verb category and time to contact for each RoI feature. The verb prediction performance is better than the baseline method, but for time to contact, using the RoI feature directly for prediction has a trivial improvement. After analysis, we found that since the convolutional neural network itself has position prior information, which is very beneficial for the prediction of time to contact depending heavily on the position, so the RoI features extracted by the SlowFast framework are more suitable.
To improve the prediction performance of time to contact, we first tried to use the boxes directly for prediction. Considering that there are clip and flip operations in data augmentation, they will change the position of boxes, so we use the raw boxes in the prediction. We normalize and input them to the MLP , and the prediction results are obtained. However, after testing, we observed that it is still not as good as the baseline method. This shows that it is not possible to use only the position information. We consider fusing the position information with the extracted features. Using the position encoding operation of the transformer, we first encode the position of the box by cosine position encoding and encode the position of the box. Then we add it to the corresponding RoI feature to complete the fusion. Such a fusion is reasonable, which is validated in Section 8.3.
2.3 Box Processing and Result Fusion
Since the generated boxes in the first stage are redundant, we need to further adjust the box number before fine-tuning the classification model. Considering that the evaluation metric is the boxes with the top-5 scores, we filtered out the top 5 boxes using the noun classification score. This operation greatly reduces the number of boxes and speeds up the inference.
Using the filtered boxes for prediction, the results on the validation set have all surpassed the baseline. To fully use existing boxes and prediction results, we fuse the prediction results using our boxes and the official predictions. The specific fusion method is to directly splice the prediction results of the corresponding keyframes in the two result files. This will introduce both boxes that are relatively similar to our box, as well as some less effective boxes. The score of the box with poor effects is usually relatively low, and it will be automatically filtered out in the evaluation program, and similar boxes are redundant. Here, NMS is used again, and we set a higher IoU threshold to remove redundant boxes.
3 Experiments
Implementation Details. We use VideoMAE-L as our backbone. The model is trained for 10 epochs using the AdamW optimizer and cosine learning rate schedule. In the fine-tuning stage, we use 8 GPUs for training, with a total batch size of 64, weight decay of 0.05, and a learning rate of for 30 epochs.
Results. We first reproduced the baseline according to the configurations and codes provided by the baseline. To improve the performance, we migrated the task to the codebase of VideoMAE . The models used in the experiments described below are all trained from this codebase. Through experiments, we found that using the transformer network and introducing the positional encoding of the boxes for training is effective for this task.
As shown in Table 8, when filtering the boxes, we not only took the top 5 but also tried to keep the top 3 and top 10 boxes, respectively. Although keeping the top-3 boxes on the validation set gave the best results, it performed mediocre on the test set. It shows that there is a certain gap between the data of the test set and the validation set, and some of our additional operations overfit the validation set. For the result fusion, we found that when the results of the top 10 boxes are selected for fusion, it produces a better effect, but it does not work when the top 3 boxes are used for fusion.
Concluding Remarks
We have presented our solutions to five tracks in the Ego4D ECCV2022 Challenge. We find a strong video backbone can give an advantage to task performance, and video backbones with various structures complementary to each other in these tasks. The typical video dataset has a distinct domain gap from the egocentric one, leading to degraded performance. We close this gap by finetuning the pre-trained video backbones on the Ego4D dataset.
Our feature aggregation method is a bit naive as we directly concatenate features from different backbones. The good craft of feature alignment or other dynamic fusion modules could benefit the egocentric video representation. Besides, adapting the employed backbone with task heads to each track is tedious. How to achieve satisfying task performance while minimizing downstream tweaking remains open.