Egocentric Video-Language Pretraining

Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Wenzhe Zhao, Weijie Kong, Chengfei Cai, Hongfa Wang, Dima Damen, Bernard Ghanem, Wei Liu, Mike Zheng Shou

Introduction

With the recent interest boom in computer vision and natural language processing, Video-Language Pretraining (VLP) has prevailed, which aims to learn strong and transferable video-language representation for powering a broad spectrum of video-text downstream tasks, such as video-text retrieval xu2016msr ; patrick2020support ; bain2021frozen , video question answering msrvttqamsvdqa ; yu2018joint ; zhu2020actbert , and video captioning krishna2017dense ; wang2018reconstruction ; zhou2018end . The success of VLP mainly stems from the availability of large-scale open-world video-text datasets miech2019howto100m , which subsume a large number of videos sourced from the Web (e.g., YouTube) and pair videos with associated textual information. For instance, HowTo100M miech2019howto100m collects 134134K hours of instructional videos accompanied by noisy narrations yielded from Automatic Speech Recognition (ASR). WebVid-2M bain2021frozen scrapes 2.52.5M descriptive videos with well-formed long captions.

Despite reaching an impressive data scale, videos in those existing video-text pretraining datasets are often of 3rd-person views and may have been edited before posting on the Web. Yet, there is a noticeable domain gap between the existing video-text pretraining datasets and 1st-person view videos such as those videos captured by wearable cameras or smart glasses. Egocentric video has received increasing interests from the academia (e.g., activity recognition caba2015activitynet , activity anticipation abu2018will , and video summarization ma2002user ) and industry (various applications in robotics and augmented reality). However, due to such a domain gap, directly transferring the existing VLP models to egocentric downstream tasks cannot fully unleash the potential of large-scale pretraining approaches, which we have confirmed in the later experimental section. To bridge this gap, we are motivated to develop Egocentric VLP models, which can greatly benefit various egocentric video downstream applications.

However, existing egocentric video datasets are of small scale and domain-specific, making Egocentric VLP prohibitive. As illustrated in Tab. 1, the formerly largest egocentric video dataset EPIC-KITCHENS-100 kazakos2019epic focuses on kitchens scenarios and its size is far smaller than those of the 3rd-person pretraining sets WebVid-2M bain2021frozen and HowTo100M miech2019howto100m . Fortunately, with the recent introduction of the massive-scale egocentric video dataset Ego4D grauman2021ego4d , it becomes possible to unlock Egocentric VLP. Ego4D consists of 3,6703,670 hours of videos with manually annotated narrations from 7474 worldwide locations, covering a large variety of daily-life scenarios and activities.

In this work, roused by the favorable scale and diversity of Ego4D, we make a significant effort to pave the way for Egocentric VLP with the following steps: (i) To address the aforementioned issue of lacking a suitable large-scale egocentric video-language pretraining dataset, we create a video-text pretraining dataset EgoClip which contains a total of 3.83.8M clean 1st-person clip-text pairs selected from Ego4D and covers diverse human daily activities. (ii) To make full use of EgoClip for video-text representation learning, we propose a novel video-text contrastive objective EgoNCE to address unique challenges in egocentric pretraining datasets. (iii) We create a development benchmark i.e., Egocentric Multiple-Choices-Question, dubbed EgoMCQ, which contains 3939K questions created from Ego4D and focuses on evaluating video-text alignment. In contrast to other downstream benchmarks, EgoMCQ has a less discrepancy from EgoClip, powering us to accurately validate and quickly iterate our designs of EgoClip and EgoNCE. (iv) We conduct extensive experiments to demonstrate the superiority of Egocentric VLP by transferring our pretrained representation to five egocentric downstream benchmarks and achieving state-of-the-art performance: 59.4%59.4\% nDCG on video-text retrieval of EPIC-KITCHENS-100 kazakos2019epic Egocentric VLP won championship on Multi-Instance Retrieval, EPIC-Kitchens Challenges @ CVPR 2022., 32.1%32.1\% mAP on action recognition of Charades-Ego sigurdsson2018charades , and significant boosts over three Ego4D challenges Egocentric VLP won championship on OSCC and 2nd place on NLQ, Ego4D Challenges @ CVPR 2022.: natural language query, moment query and object state change classification.

Related Work

Video-Language Pretraining. The introduction of large-scale video-text datasets miech2019howto100m ; bain2021frozen has enabled the emergence of VLP approaches to improve the video-text representation for various vision-language tasks anne2017localizing ; chen2017sca ; msrvttqamsvdqa , such as MIL-NCE which miech2020end proposes to match clips with multiple captions close in temporal to adapt the video-text misalignment of HowTo100M miech2019howto100m . Dominant VLP methods can be classified into two groups, namely: joint- and dual-encoders. The former combines videos and texts as a single input to the encoder that performs the multimodal fusion. For instance, lei2021less ; Sun_2019_ICCV concatenate videos and texts together before feeding them to a unified transformer. Conversely, methods like bain2021frozen ; wang2022object exploit dual encoders to independently project the video and text inputs into a common space and minimize the distance between the paired representations. These approaches are preferred in retrieval settings as they allow for efficient indexing of a single modality escorcia2019temporal ; miech2021thinking . For example, Frozen bain2021frozen employs two separate transformers to encode video and text features and aligns them by video-text InfoNCE infonce . In our work, we adopt the Frozen bain2021frozen but extend its InfoNCE to EgoNCE via positive and negative sampling for egocentric-friendly pretraining.

Egocentric Video Datasets. Egocentric videos, collected by participants using wearable cameras, offer a natural perspective of people’s daily activities and raise a range of challenging research topics caba2015activitynet ; abu2018will ; wong2022assistq . Several egocentric video datasets have been developed in decades, e.g., damen2022rescaling ; sigurdsson2018charades ; li2015delving . However, since the collection of egocentric videos is expensive, previous egocentric datasets tend to be small-scale and domain-specific. These limitations hinder 1st-person view research and fail to match the progress of 3rd-person counterparts, such as VLP miech2020end ; lei2021less ; bain2021frozen . Recently, a massive egocentric video dataset Ego4D grauman2021ego4d has been released, which consists of 3,6703,670 hours of videos collected by 931931 people from 7474 worldwide locations in 99 different countries, where most videos are accompanied by narrations, audio, 3D meshes, and more. Furthermore, Ego4D introduces a suite of new challenging benchmarks (e.g., Natural language query and moment query) to fully explore the 1st-person visual experience. With this step-changing dataset and benchmarks, Ego4D would lead to a new research surge on egocentric visual perception.

EgoClip: An Egocentric Video-Language Pretraining Dataset

Data curation. For our EgoClip dataset, we source data from Ego4D grauman2021ego4d , which contains 9,6459,645 untrimmed videos of varying lengths from 55 sec to 77 hrs. From these videos, most are associated with dense timestamp-level narrations assigned by two different annotators, describing the camera wearer’s activities and interactions with objects. For example, the narration “#C C puts the scrapper down.” corresponds to video content that occurred at 3.70s3.70s, where “#C” refers to the camera-wearer. Notably, narrations in Ego4D are well-aligned with the videos, both temporally and visually. Prior pretraining datasets are characterized by a much greater level of temporal misalignment between the video and text (e.g., HowTo100M miech2019howto100m narrations are scraped from ASR, yielding sentences misaligned or even unrelated to video content). We first filter Ego4D videos with missing narrations (7.4%7.4\% of the total video duration) and exclude videos that belong to the validation and test sets of the Ego4D benchmark challenge grauman2021ego4d (a further 23.9%23.9\% of the total video duration). Next, we retain textual annotation from both narrators in EgoClip, allowing us to consider narration diversity when pairing video and text for pretraining purposes. Finally, we adopt several criteria to filter the video and textual narrations, further reducing noise (detailed steps are provided in Supplementary B.1). Overall, this procedure yields 2.92.9K hours of videos with 3.853.85 million narrations which cover 29272927 hours of video from 129129 different scenarios. EgoClip has 21.921.9 clips per minute with an average clip length of 1.01.0 seconds and a standard deviation of 0.90.9 seconds (the longest clip is up to 6060s). Additional analyses are included in the Supplementary B.3.

Creation of clip-text pairs. Clip-text pairs are the common data format for VLP, but are usually not present in untrimmed video datasets with only a weak matching between narrations captions and videos. This was first discussed in HowTo100M miech2019howto100m , which pairs subtitles to video clips with corresponding time intervals to produce noisy pairs. This is not suitable for Ego4D since each narration is annotated with a single timestamp rather than an interval. Thus, we design a contextual variable-length clip pairing strategy. Formally, narrations per video in Ego4D are organized as a sequence of sentences {T0,⋯ ,Tn}\{\mathcal{T}_{0},\cdots,\mathcal{T}_{n}\} with exact timestamps {t0,⋯ ,tn}\{t_{0},\cdots,t_{n}\}, indicating an event ii described by Ti\mathcal{T}_{i} happened in the moment tit_{i}. For a narration Ti\mathcal{T}_{i} with timestamp tit_{i}, we pair a clip Vi\mathcal{V}_{i} with following start and end timepoints:

which represents a window centered around the timestamp tit_{i} with temporal duration equal to βi/α\beta_{i}/\alpha. βi\beta_{i} is an adjustable parameter equal to the average temporal distance between pairs of consecutive narrations, i.e., ∑j=0n−1(tj+1−tj)/n{\sum_{j=0}^{n-1}(t_{j+1}-t_{j})}/{n}. We compute βi\beta_{i} on a per video basis. Conversely, α\alpha is a scale factor computed as the average of all βi\beta_{i} across all videos in the EgoClip (α=4.9\alpha=4.9 seconds). Intuitively, Eq. 1 is derived from three observations: (i) Centering tit_{i} helps involve prior information about the event ii; (ii) βi\beta_{i} measures the clip duration according to its scenario, such as longer clips watching television (352.9352.9 seconds) v.s. shorter clips harvesting crops (0.90.9 seconds); (iii) α\alpha controls the context granularity of clips (e.g., a large α\alpha pays more attention to rapid, atomic actions). We ablate these design choices in our experimental section.

Video-Language Pretraining Model

To efficiently transfer video-language representation to egocentric downstream tasks (e.g., video-text retrieval on EPIC-KITCHENS-100 damen2022rescaling ), We prefer the dual-encoder (discussed in Sec. 2) as our VLP model architecture. In particular, we emphasize devising a general pretraining objective EgoNCE to adapt the existing VLP model to the egocentric domain (e.g., EgoClip).

We choose Frozen bain2021frozen as our pretraining architecture. Frozen bain2021frozen design encompasses an elegant and simple dual encoder strategy (one per modality) which has favorable characteristics (e.g., indexability and efficiency escorcia2019temporal ; miech2021thinking ). Note that this allows us to use our pretrained network in single-modality tasks (e.g., video-only tasks). In practice, the video encoder adopts the TimeSformer timesformer architecture, while the text encoder builds upon DistillBERT distilbert . However, our approach is not limited to the encoder’s design (e.g., the video backbone can be replaced by SlowFast slowfast or Video Swin liu2022video ). In the rest of the paper we adopt this notation: (Vi,Ti)(\mathcal{V}_{i},\mathcal{T}_{i}) represents the video-text input to the model, while vi\mathbf{v}_{i} and ti\mathbf{t}_{i} are used to identify the video and text embeddings.

2 EgoNCE: An Egocentric-friendly Pretraining Objective

A common pretraining objective for the dual-encoder VLP is InfoNCE infonce , where the matching visual-text pairs in the batch are treated as positives while all other pairwise combinations in the batch are regarded as negatives. Formally, within a batch B={1,⋯ ,N}\mathcal{B}=\{1,\cdots,N\}, InfoNCE is computed by the sum of the video-to-text loss Lv2t\mathcal{L}_{\text{v2t}} and text-to-video loss Lt2v\mathcal{L}_{\text{t2v}}. For simplicity, we only formulate Lv2t\mathcal{L}_{\text{v2t}}, whereas Lt2v\mathcal{L}_{\text{t2v}} is defined in a symmetric way:

where the ii-th video embedding vi\mathbf{v}_{i} and jj-th text embedding tj\mathbf{t}_{j} are L2L_{2} normalized features, and τ\tau is a temperature factor.

However, this simple objective performs not well on large-scale video-text datasets like HowTo100M miech2019howto100m due to the serious misalignment between the two modalities of data. Therefore, mil_nce proposes MIL-NCE which treats temporal nearest captions as positive samples.

In this work, our 1st-person human daily activity dataset, i.e. EgoClip, presents two unique challenges compared to the existing 3rd-person view video-text datasets: Challenge (i): The same action often occurs in different scenarios (e.g., “unlock the phone” could happen when “lying in bed” or “walking outdoors”). Challenge (ii): Often, different actions appearing in the same scenario tend to have indistinguishable visual differences (e.g., when “working in front of the laptop”, “typing on the keyboard” or “moving the mouse” have similar feature representations).

To overcome these two unique challenges, we propose a novel EgoNCE training objective which takes into account two simple yet efficient sampling strategies based on the vanilla InfoNCE.

Action-aware Positive Sampling. In this work, we make a reasonable assumption that the critical elements in linking visual actions to textual narrations are verbs and objects mentioned in the narrations (e.g., “drinking coffee” and “opening fridge”). Following this assumption, we can devise a clever method to address challenge (i). Specifically, for each narration, we identify its nouns and verbs and merge synonym words based on the Ego4D taxonomy dictionary grauman2021ego4d , a thesaurus recording meaningful nouns/verbs in Ego4D narrations. Then, batch samples that shared at least one noun and at least one verb are treated as positive samples. At last, for the sample ii, we define its positive samples set within batch B\mathcal{B} as Pi={j∈B ∣ noun(j)∩noun(i)≠∅,verb(j)∩verb(i)≠∅}\mathcal{P}_{i}=\{j\in\mathcal{B}~{}|~{}\text{noun}(j)\cap\text{noun}(i)\neq\varnothing,\text{verb}(j)\cap\text{verb}(i)\neq\varnothing\}.

Scene-aware Negative Sampling. To address challenge (ii), we consider different actions in the same scenario as hard negative samples. Specifically, for each video clip ii, we sample an adjacent clip i′∈N(i)i^{\prime}\in\mathcal{N}(i), which is close to ii in time within the same video. We augment the original batch B\mathcal{B} with such hard negative samples and each sample ii in B\mathcal{B} has its negative counterparts i′i^{\prime}. Hence the batch is updated as B~={1,2,⋯N⏟B,1′,2′,⋯ ,N′⏟N(B)}\mathcal{\widetilde{B}}=\{\underbrace{1,2,\cdots N}_{\mathcal{B}},\underbrace{1^{\prime},2^{\prime},\cdots,N^{\prime}}_{\mathcal{N}(\mathcal{B})}\}.

With these two sampling strategies, our new pretraining objective EgoNCE can be formulated as:

Here the item in purple corresponds to our proposed action-aware positive samples and blue corresponds to our proposed scene-aware negative samples. EgoNCE provides a general extension to adapt the existing VLP models for video-text pretraining datasets in the egocentric domain.

EgoMCQ: A Benchmark for Egocentric VLP Development

The need for a development benchmark. We find that most egocentric benchmarks are domain-specific and focus on single-modality tasks (see Tab. 1). However, our purpose is to exploit Ego4D’s diversity to learn rich video-text representations. Hence, to validate our design choices of the pretraining dataset (e.g., EgoClip), and model (e.g., EgoNCE), it is essential to measure performance on a benchmark highly aligned with the pretraining task. Therefore, we propose EgoMCQ, a new egocentric benchmark for reliable and fast developments of Egocentric VLP.

Data source. We start from the Ego4D data excluded from constructing the EgoClip, which mainly covers the validation set of the Ego4D challenge benchmarks. Additionally, to assure that the scene is not visible during pretraining, we manually remove videos that share multiple views with the videos in EgoClip. To ensure diversity, we randomly select one annotator’s narration for each video. We follow the same clip pairing strategy as Eq. 1 to be consistent with the data format of EgoClip.

Benchmarking task design. To determine the task for development, we first consider video-text retrieval since it highly aligns with the VLP pretraining objective. However, as depicted in the top half of Fig. 2, for an action (e.g., close the refrigerator), there are substantial duplicates or semantically similar captions in Ego4D. This can cause issues in retrieval evaluation wray2021semantic making model training unreliable. A straightforward approach to prevent this is deduplication (dedup), but it is challenging to devise a dedup criterion and perform well in the retrieval settings of a “one-to-whole validation set”. Therefore, we select the Multiple-Choice Questions (MCQ) task for development since repetitions are highly unlikely given a small number of answers.

Grouping strategies. To set up the MCQ task, a naive construction randomly groups five video clips to form options for a question. But we find randomly grouping is not challenging since options are highly likely to come from different videos and vary widely in content. We redefine this basic setting as “inter-video” and ensure that the five clips originate from different videos, aiming to distinguish instances from different scenarios (the left-bottom of Fig. 2). Furthermore, we propose a more challenging setting “intra-video” by grouping five continuous clips together.This setting is regarded as a specific form of video-text localization focused on fine-grained context clues, such as hand interaction (the right-bottom of Fig. 2). Dedup is performed within five options for each question for reliable assessment (see Supp. C.1) and we adopt accuracy as the EgoMCQ metric.

Statistics. We finalize 3939K questions covering 198198K narrations with 468468 hours of video, where the “inter-video” has 2424K questions covering 290.3290.3 hours of videos. And the “intra-video” has 1515K questions and covers 178.3178.3 hours of videos. The average duration among the five options is 34.234.2 seconds (More statistics of EgoMCQ are shown in Supplementary C.3).

Experiments

We assess our Egocentric VLP along two directions: (i) We conduct an extensive analysis to explore key components of Egocentric VLP (e.g., EgoClip, EgoNCE, and EgoMCQ); (ii) we transfer our pretrained model to various downstream tasks to validate the quality of our video-text representation.

We evaluate our VLP model on five egocentric benchmarks, spanning video-text tasks and pure video tasks, across three different datasets. We briefly describe each task below.

Multi-Instance Retrieval of EPIC-KITCHENS-100. This task is modelled as a video-text retrieval which considers the semantic overlap between different videos narrations, where multiple videos may correspond to the same narration. The training set contains 67.267.2K clips and validation set contains 9.79.7K clips. The evaluation metrics are mean Average Precision (mAP) and the normalized Discounted Cumulative Gain (nDCG).

Natural Language Query of Ego4D Challenges. The Natural Language Query task is modelled as a natural language grounding problem hendricks2018localizing ; Gao_2017_ICCV ; soldan2021vlg . Given a language query and a video, the task aims at localizing the temporal interval within the video, in which the answer is deducible. The training set contains 11.311.3K queries annotated from 11K clips for this task, while the validation contains 3.93.9K queries collected from 0.30.3K clips. The evaluation metric is Recall@KK for IoU=θ{=}\theta (R@KK-IoU=θ{=}\theta) hendricks2018localizing where θ\theta is a threshold. We evaluate for K∈{1,5}K{\in}\{1,5\} and θ∈{0.3,0.5}\theta{\in}\{0.3,0.5\}.

Action Recognition of Charades-Ego. This dataset has 6464K instances, spanning 1st-person and 3rd-person views and covering 157157 activity categories for training. We train and evaluate only on the 1st-person videos. The validation set contains 847847 videos for classification and each video belongs to multiple classes. The evaluation metric is mAP.

Moment Query of Ego4D Challenges. The Moment Query task is a video-only task modelled as Temporal Action Localization caba2015activitynet . Given a particular high-level activity category, the task solution consists of retrieving all the possible temporal windows where the activity occurs. The training set contains 13.613.6K instances from 1.51.5K clips, while the validation set contains 4.34.3K instances from 0.50.5K clips. The evaluation metrics are mAP and R@KK-IoU=θ{=}\theta for K∈{1,5}K{\in}\{1,5\} and θ∈{0.3,0.5,0.7}\theta{\in}\{0.3,0.5,0.7\}.

Object State Change Classification (OSCC) of Ego4D Challenges. This OSCC task is modelled as an (N+1)-way classification aiming to identify an object’s state change in a given video. The training and val. sets contain 4141K and 2828K clips, respectively. The evaluation metric is accuracy.

Implementation Details. Our codebase is based on the official Frozen https://github.com/m-bain/frozen-in-time one and retains the same settings unless specified. During pretraining, we sample 44 frames for each clip, and use the Adam optimizer kingma2014adam with a learning rate of 3×10−53{\times}10^{-5}. To select the best method we pretrain our architecture for 1010 epochs and use the best performing model on the EgoMCQ benchmark. Pretraining takes two days on 3232 A100 GPUs (1,5361,536 GPU hrs).

2 Ablation Studies

We validate our proposed strategies, i.e., Eq.1 in Tab. 2, by comparing the following variants: (a) fixed length α\alpha, start at timestamp; (b) fixed length α\alpha, center at timestamp; (c) variable clip, start and end by adjacent timestamps; (d) our proposed strategy, scaled by 22; (e) our proposed strategy, scaled by 44; (f) our proposed strategy.

We consider that a good pretraining dataset creation strategy should satisfy: (1) the VLP model trained on EgoClip should be able to well distinguish instances in EgoMCQ with the same data format; (2) the VLP model pretrained on EgoClip with the specific clip creation strategy should perform well on public downstream tasks (e.g., video-text retrieval on damen2022rescaling and zero-shot for efficiency).

We draw several conclusions from Tab. 2: (i) The performance of EgoMCQ is well aligned with the zero-shot result on EPIC-KITCHENS-100, especially minor gain on downstream but noticeable on EgoMCQ, which means EgoMCQ provides valid feedback and is suitable as a development set. (ii) Under the same clip length α\alpha, (b) surpassing (a) proves that centering at timestamp includes prior information is helpful. (iii) Variable-length clips make a big difference, as shown in (c) and (d).

Notably, with our designed βi\beta_{i}, (d) outperforms (b) with a similar average clip length, which validates our key idea of “contextual varied clip length”. (iv) Based on (d), (e), and (f), we found a proper scale factor greater than 11 is preferred, which helps focus on a large of instantaneous actions densely labeled by Ego4D grauman2021ego4d . These ablation studies demonstrate the effectiveness of our proposed EgoClip creation strategy and EgoMCQ for development.

Effect of EgoNCE. In this section, we evaluate the effect of the proposed sampling strategies for the EgoNCE objective (Eq. 3) on EgoMCQ and compare against a vanilla InfoNCE loss (Eq. 2). We ablate several configurations for positive and negative sampling strategies. The sampling strategy for positive pairs exploits language cues, while negative pairs rely on temporal, visual cues. Given a text-video pair, we regard other text-video pairs as positive if the textual narrations: (a) share at least one noun, (b) share at least one verb, and (c) share at least a verb-noun pair. Conversely, we define the following heuristics for negative sampling: (d) a random text-video pair from EgoClip, (e) a text-video pair from the same video, and (f) a text-video pair within 11 minute from the given video-text pair annotation timestamp. Tab. 3 shows that using solely verbs (a) or nouns (b) for positive selection degrades the accuracy performance with respect to naive InfoNCE. However, we successfully push the performance beyond the baseline results when considering both verbs and nouns jointly (c). Moreover, we notice that merely selecting negatives within the same video leads to better performance. In particular, we obtain the best performance for temporally “hard negatives” (f). Finally, we pick the optimal settings from positive and negative sides and combine them together for (g) EgoNCE and reach the best results.

3 Comparisons with State-of-the-arts

Multi-Instance Retrieval. In Tab. 4, we report both zero-shot and fine-tuning evaluation results. In the zero-shot setting, pretraining with EgoClip (3.83.8M), despite being smaller in scale, still outperforms CC3M+WebVid-2M (5.55.5M) and HowTo100M (136136M), validating the unique benefit of pretraining on egocentric data. When fine-tuned with 44 frames (rows 5-9), EgoClip pretraining maintains a margin over the best baseline CC3M+WebVid-2M, further verifying the viewpoint domain gap within fine-tuning. Lastly, we increase the sample frames of our finalized model as well as the best competitor CC3M+WebVid-2M pretraining to 1616 (rows 10-11). As expected, performance gains accompany the frame increase. We deem that notable benefits come from better temporal modeling for frequent action in the 1st-person view. Overall, our pretraining model outperforms the best baseline (JPoSE) by 1.01.0 mAP and 5.9%5.9\% nDCG while requiring fewer frames and input modalities.

Natural Language Query. We report validation results on Tab. 5. We adopt the same baselines as introduced in grauman2021ego4d , namely: 2DTAN zhang2020learning and VSLNet zhang2020span , and substitute the SlowFast-BERT features with our video and language representations. We observe a large boost in performance offered by our pretrained model on all metrics. Notably, we improve R@11 for IoU=0.30.3 from 5.455.45 to 10.8410.84, despite our video branch not being pre-trained on Kinetics400. Besides, we significantly surpass VLP pretrained on CC3M+WebVid-2M and HowTo100M. We believe that this increase is due to the egocentric data availability and the video-text interaction learned from large-scale pretraining. Please see Supplementary E.5 for the test set results.

Action Recognition. We conduct action recognition on Charades-Ego, where categories are short phrases like “Holding some clothes”. Thus this task can be solved as a video-text retrieval by leveraging the text representation. We present the result in Tab. 6 under zero-shot and fine-tuning settings. In zero-shot settings, our model outperforms two supervised baselines, which validates the stronger generalization of jointly learning video-text features. After fine-tuning (rows 5-9), our model surpasses all VLP counterparts and improves over the state-of-the-art classifier Ego-Exo by 2.02.0% with fewer sampled frames, which shows the superior advantage of joint video-text representations.

Moment Query. This task investigates the quality of video-only features. We extract video features and provide them as input to the VSGN model zhao2021video . We report the validation results in Tab. 7, We find that our features achieves the best performance over SlowFast features with an increase of 4.66%4.66\% in Avg mAP. Moreover, we maintain better performance with respect to 3rd-person large-scale pretraining datasets. This demonstrates that the 1st-person VLP model also learns competitive video representations. Please see the Supplementary E.6 for the test set results.

Object State Change Classification. We report the validation results on Tab. 8. Once again, our model achieves the best performance of all baselines, 2.4%2.4\% than CC3M+WebVid-2M counterparts, which indicates our visual representations are able to focus on the fine-grained clues related to state changes.

Summary of EgoNCE. From the above experimental results, Frozen pretrained on EgoClip with the EgoNCE objective brings a consistent improvement over the InfoNCE on all downstream tasks, which comprehensively demonstrates the effect of EgoNCE, as well as the decision from EgoMCQ.

Conclusion, Limitations, and Societal Impacts.

To the best of our knowledge, this work is the pioneering work to unlock Egocentric VLP. (i) We devise a principled data curation and create EgoClip, an egocentric large-scale text-video pretraining dataset with 3.83.8M clip-text pairs well-chosen from Ego4D. (ii) We exploit the particular characteristics of egocentric videos and devise EgoNCE with meaningful sampling strategies for effective egocentric pretraining. (iii) We create EgoMCQ, an egocentric video-language benchmark close to the pretraining set to support efficient exploration and development of EgoClip and EgoNCE. Finally, we further demonstrate the strong representation of our egocentric pretraining on five tasks across three datasets. We believe that our EgoClip, EgoMCQ and EgoNCE would greatly benefit the egocentric video community, laying a good foundation for the new research trend of egocentric VLP. Limitations. Our pretraining approach does not take into account the long-term temporal dependencies in long Ego4D videos. We leave this for future work. Societal impact. Egocentric VLP learns real-world perception knowledge that may contribute to practical applications such as augmented reality and robotics. However, Ego4D videos collected by participants may contain users’ privacy and unintended biases, so should be used cautiously. We refer the readers to the Ego4D paper about further privacy and societal impacts.

Acknowledgements

This project is supported by the National Research Foundation, Singapore under its NRFF Award NRF-NRFF13-2021-0008, and Mike Zheng Shou’s Start-Up Grant from NUS. The computational work for this article was partially performed on resources of the National Supercomputing Centre, Singapore. Michael Wray and Dima Damen are supported by EPSRC UMPIRE (EP/T004991/1). Mattia Soldan and Bernard Ghanem are supported by the King Abdullah University of Science and Technology (KAUST) Office of Sponsored Research through the Visual Computing Center (VCC) funding, as well as, the SDAIA-KAUST Center of Excellence in Data Science and Artificial Intelligence (SDAIA-KAUST AI). Thanks to Tencent Data Platform for the support of computing resources. Our work is built upon the Ego4D dataset, and we greatly appreciate the contributions and efforts of the Ego4D community.

References

Appendix

We present the following items in the supplemental material:

Differentiating Egocentric VLP and Ego4D in Sec. A.

Construction details and statistics of EgoClip pretraining dataset in Sec. B.

Construction details and statistics of EgoMCQ benchmark in Sec. C.

Technical details of our VLP model in Sec. D.

Additional experimental details and results in Sec. E.

A Differentiating Egocentric VLP and Ego4D

In our work, we study the video-language pretraining in a specific yet significant domain - the 1st-person view, which is motivated by the release of the Ego4D dataset. However, there is a long way to pave from the Ego4D dataset to Egocentric VLP, which consists of the pretraining dataset, development set, model designs, and transferability evaluation. Since they are not as fully explored as their third-person counterparts, thus we pioneer them by ourselves and conduct a systematic study toward the egocentric video-language pretraining - the contribution of our work.

Despite the merits of Ego4D, it has not been proposed for video-language pretraining, and cannot be directly used as its untrimmed videos, no direct video-text pairs, and noisy data. We thus see our clear distinction and contribution in proposing a successful approach to curate a pretraining dataset, our proposed EgoClip. Notably, It is also non-trivial to figure out what is the best way of curating Ego4D to create a pretraining dataset EgoClip, e.g., our pairing approach outperforms the naive strategy with a large margin in the development set, which requires substantial design and experimental validations. We add a Tab. 9, as an extension of Tab. 1, to clearly show their difference.

A.2 Development set

In the 1st-person domain, there is lacking a satisfactory benchmark that good aligns with pretraining data diversity and focuses on video-text alignment. Therefore, we propose a new development set i.e. EgoMCQ to power rapid design of video-text pretraining i.e. its pretraining dataset and model pretraining objective.

A.3 Model designs

We select Frozen as the baseline because its elegant and scalable dual-encoder architecture is representative in state-of-the-art VLP methods. Besides, corresponding to MIL-NCE built on top of the 3rd-person domain’s HowTo100M , we aim to explore a general pretraining objective i.e., EgoNCE to learn rich video-text representations in 1st-person domains.

A.4 Transferability evaluation

Extensive experiments and promising results demonstrate the effectiveness and necessity of Egocentric VLP, which will greatly benefit the egocentric community. Note that Ego4D has not been used previously for any downstream tasks on other datasets. This is also where our work makes significant value.

B Construction details and statistics of EgoClip pretraining dataset

After we source video-text data for EgoClip, we adopt the following criteria to further reduce noise:

(i) We select double-sized stereo videos (1.3%1.3\% videos dur) and keep half per video for a normal size.

(ii) We discard videos with an aspect ratio greater than 22 (0.4%0.4\% videos dur).

(iii) We filter narrations with unsure tags (4.0%4.0\% texts) e.g. “#C C washes #unsure in sink”.

(iv) We remove narrations less than 33 words (0.9%0.9\% texts), since such narrations generally cannot be deduced from the video, e.g., “#C C speaks”, “#C C looks”.

B.2 Data compression

The Ego4D videos are untrimmed, which tend to be very long (average 2424 mins and max to 77 hrs) and have large resolution (e.g., 1920×10801920\times 1080, 1440×10801440\times 1080), so it is impossible to adopt untrimmed videos as model input due to heavy data loading. Therefore we propose to compress them:

(i) We first resize all videos with short size 256256.

(ii) Chunk each resized video into several segments, which are up to 1010 min in length.

During pretraining, given the start and end time points of a clip, we only load the segment that this clip belongs to, rather than the whole video. To this end, we are able to perform efficient end-to-end pretraining with raw RGB videos as model input. One epoch of pretraining 3.8M3.8\text{M} video-text pairs costs 66 hrs on 3232 V100 GPUs (192192 GPU hrs).

B.3 Data analysis

Geographic diversity. We present the distribution of EgoClip clips source in Fig. 3, which covers worldwide 13 institutions from 9 different countries , including: Europe (UK, Italy); Asia (India, Japan, Singapore, Kingdom of Saudi Arabia); America (USA, Colombia); Africa (Rwanda). Therefore, our created pretraining dataset inherited the good geographic as well as participants diversities of Ego4D (More details can be found in “Supp. C. Demographics” in Ego4D paper ).

Scenario diversity. We have statistics the scenario distribution of EgoClip in Fig. 5, which covers 129129 human daily scenarios e.g., household (cooking, cleaning), outdoor (shopping, hiking), workplace (at desk, on a laptop), leisure (playing board games), etc. Notably, this distribution is long tailed, where the largest scenario “Crafting/knitting/sewing/drawing/painting” includes 622K (11.1%)622\text{K}~{}(11.1\%) and the smallest scenario “Hair and Makeup stylist” contains 3535 instances.

Clip analysis. We present the statistics on the created clips in EgoClip. Fig. 6 (a) shows the distribution of clip frequency over the 2.9K2.9\text{K} pretraining set videos (For each video, we calculate two frequencies from two annotators respectively). The varying clip frequencies are mainly dependent on manual narrations that are annotated based on the video scenarios and activities. There have average 13.413.4 clips per minute of video, maximize to 175.8175.8 narrations / minute and minimize to 0.060.06 narrations / minute. Our clip creation strategy Eq. (1) takes this characteristic into account by estimating clip length based on the frequency of the video that the clip belongs. Fig. 6 (b) displays the distribution of clip duration. The average duration is 0.980.98 seconds with a standard deviation of 0.950.95 seconds, and 69.5%69.5\% of clips are less than 1.01.0 seconds in length, due to the massive atomic instantaneous actions densely labeled by Ego4D. Besides, the clip might be max to 65.3665.36 seconds, which corresponding to the scenario that “a people walking in a forest”.

Narration analysis. In Fig. 6 (c), we present the distribution of narration words length. The average words length of EgoClip narration is 9.399.39. Notably, the EgoClip narrations cover 116116 verbs and 555555 nouns, where we merge the semantically synonyms words, e.g., the nouns of “handkerchief”,“napkin”,“serviette”,“tissue”,“wipe” both belong to “napkin”. Each narration of EgoClip have 1.841.84 nouns and 0.870.87 verbs on average.

We further display the distribution of the top 50 most frequently verbs and nouns of EgoClip in Fig. 7. The most common nouns is “napkin”, which appeared in 1.0M (27.06%)1.0\text{M}~{}(27.06\%) clips.

Visualizations. In Fig. 8, we visualize some clip-text pairs created by our strategy.

C Construction details and statistics of EgoMCQ benchmark

To ensure repetitions do not appear in five options, we devise a deduplication strategy. Initially, we use Sentence-BERT to extract sentence-level embeddings of narrations and set a manual threshold to remove repetitions. But in this way, it is hard to control the fine-grained diversity between narrations, e.g., two narrations “#C C closes the refrigerator with his left hand.” and “#C C opens the refrigerator with his left hand.” only differ in one word. These two sentences have a high score in sentence-level similarity, but are entirely different in semantic meanings. We hope to keep them and let the model distinguish them, especially in our intra-video setting.

Therefore, we propose to extract the first verb and the first noun of each narration and use them to define a tag for each narration. The narrations shared with the same verb and the noun will be assigned the same tag. We also consider the words synonyms (based on Ego4D taxonomy dictionary ). For instance, “#C C take the phone” and “#C C pick the cellphone” are semantically same in verb and noun thus will be assigned the same tag. Then the narrations shared with the same tag are treated as repetitions, we only keep one of them and sample a new one until the tags of the five options are different.

C.2 Multiple-views removing

We first select videos from NUS/Minnesota/Georgia Tech/Indiana sources, which contribute to the multi-camera video data. Then, based on the metadata of the video (i.e. times when videos were collected), we observed that videos collected in the same timeframe tend to be multi-views of the same recording, so we manually group these videos into the same split to ensure the same scene does not appear in another split.

C.3 Data analysis

We finalize 3939K questions covering 198198K narrations with 468468 hours of video, where the “inter-video” has 2424K questions covering 290.3290.3 hours of videos. And the “intra-video” has 1515K questions and covers 178.3178.3 hours of videos. The average duration among the five options is 34.234.2 seconds.

Geographic diversity. We present the geographic diversity of EgoMCQ in Fig. 10, which covers 13 institutions and is align with the geographic diversity of EgoClip.

Scenario diversity. In Fig. 5, we present the scenario distribution of EgoMCQ, which covers 7474 scenario. The largest scenario “Cooking” includes 49K (15.3%)49\text{K}~{}(15.3\%) clips and the smallest scenario “Bus” contains 6 instances. EgoMCQ covers 71%71\% of scenarios in EgoClipand has other 33 scenarios not appear in EgoClip. EgoMCQ is close to EgoClip both in terms of geography and scene diversity, making it a good development set for EgoClip pretraining.

EgoMCQ covers 198K198\text{K} narrations and each narration contains 3.153.15 nouns and 0.970.97 verbs in average. In Fig. 11, we display the top 50 most frequently verb and nouns of EgoMCQ. The mostly common noun is “hand”, covering 86K (36.2%)86\text{K}~{}(36.2\%) instances and the mostly frequently verb is “pick”, which covers 28K (12.0%)28\text{K}~{}(12.0\%) instances.

Visualization. In Fig. 9, we display examples of both the intra and inter settings of EgoMCQ.

D Technical details of our VLP model

In this section, we present more technical details of our VLP model, mainly architecture and pretraining objective.

D.2 Pretraining objective: EgoNCE

To supplement the Eq. 2 and Eq. 3, we first formulate the complete form of InfoNCE:

and our EgoNCE extends the above as Eq. 5 via two sampling strategies:

For positive sampling (the numerator term), we pre-extract the nouns and verbs for each narration Ti\mathcal{T}_{i} before pretraining and define two word vectors win∈{0,1}K1\mathbf{w}_{i}^{n}\in\{0,1\}^{K_{1}} and wiv∈{0,1}K2\mathbf{w}_{i}^{v}\in\{0,1\}^{K_{2}} to encode the appearing nouns and verbs in sentence, where K1K_{1} and K2K_{2} denote the number of nouns and verbs in EgoClip (Refer to Sec B narration analysis). During pretraining, for another instance jj within batch, we calculate the sij=(win)Twjn⋅(wiv)Twjvs_{ij}=(\mathbf{w}_{i}^{n})^{T}\mathbf{w}_{j}^{n}\cdot(\mathbf{w}_{i}^{v})^{T}\mathbf{w}_{j}^{v}, if sij>0s_{ij}>0, we regard instance jj is one of the positive sample j∈Pij\in\mathcal{P}_{i} of instance ii. Notably, the positive sampling space P\mathcal{P} would cover B~\mathcal{\widetilde{B}} when working with the negative sampling strategy.

For negative sampling (the denominator term), each time we sample an instance ii, we sample an instance i′∈Vii^{\prime}\in\mathcal{V}_{i} in the same video and close in time (less than 1 min) to generate the negative sample i′∈N(i)i^{\prime}\in\mathcal{N}(i) of instance ii. Notably, in this way, the actual instance within the batch ∣B~∣=2N|\mathcal{\widetilde{B}}|=2N will be double the batch size ∣B∣=N|\mathcal{B}|=N. In practice, we have to halve the batch size due to GPU memory limitations. Under halving the batch size, random sampling doesn’t help in our method, which can be concluded by comparing baseline InfoNCE and variants (d) in Tab. 3 of the main body, where the batch size of the latter is half of the former. Despite this, our proposed sampling strategy (f) can successfully improve the pretraining effect beyond baseline.

In contrast to the conventional negative sampling from the same video , we specifically design our temporally adjacent negative sampling strategy to focus on the frequent appearance changes in egocentric videos, which has not been explored in previous approaches.

E Additional experimental details and results

Following the settings of official Frozen , the video encoder is initialized with ViT weights trained on ImageNet-21K with sequence dimension D=768D=768. The text encoder is based on huggingface’s distilbert-base-uncased. The dimension of common feature space is set as 256256, and the temperature parameter is set to 0.050.05. During pretraining, each video is resized to 224×224224\times 224 as input with sample frames number 44 and batch size 512512. We use the Adam optimizer with a learning rate of 3×10−53\times 10^{-5} with a total epoch of 1010. When transferring to downstream tasks, we select the checkpoints with the best score on EgoMCQ benchmark i.e. average accuracy of inter-video and intra-video settings by default.

E.2 Downstream settings

We present the setting details of the downstream tasks we evaluated. For a fair comparison, for VLPs variants pretrained on different datasets, we use the same settings on downstream tasks, such as the fine-tuning objective.

EPIC-KITCHENS-100 Multi-Instance Retrieval. In this task, after we finalize video-text pretraining, we continue to fine-tune the VLP model and keep most settings of pretraining (e.g., input resolution, learning rate). Notably, we set the training epoch as 100100 and replace the training objective as Multi-instance Maxmargin loss in Eq. 6, which is same as the baseline method JPoSE . The reason for this is that in this task a narration may be jointly associated with multiple clips, so multi-instance learning mechanism can better handle such a situation. And this dataset also provides the action label to calculate the correlation cijc_{ij} between two clip-text pairs (i,j)(i,j), which supports the implementation of Multi-instance Maxmargin loss.

where Ω={(i,j,k) ∣j∈i+,k∈i−}\Omega=\{(i,j,k)~{}|j\in i^{+},k\in i^{-}\} is a triple, which indicates a positive instance jj and a negative instance kk for ii. In our setting, we define the positive set as i+={j∣cij>0.1}i^{+}=\{j|c_{ij}>0.1\} and the negative as the remains sample within batch. The γ\gamma is a margin factor and we set it as 0.20.2.

Charades-Ego Action Recognition. In this task, the textual categories are short phrases like “Holding some clothes”. Thus, we regard this task as a kind of video-text retrieval by leveraging the text representation and using the InfoNCE as fine-tuning objective. We set the epoch number as 1010 and keep other parameters unchanged.

Ego4D Natural Language Query This task is a kind of video-text localization and is hard to perform end-to-end training (since a clip might long to 12001200 seconds). The baseline method takes 23042304 dim SlowFast features (1.871.87 fps, with Kinetics 400 pretrained) and 768768 dim BERT features as input. Therefore, we propose to replace the baseline input features as features of pretrained VLP video and text encoders to evaluate the pretraining effectiveness. We extract the features with the same fps 1.871.87 and sampling frame number 44. In fine-tuning stage, we keep the default setting of .

Ego4D Moment Query This task is a video-only task: temporal action localization. Similar to Natural Language Query task, we replace the input Slowfast features of baseline VSGN with VLP video features for evaluation. The extraction details are the same as Natural Language Query.

Ego4D Object State Change Classification This is an action classification task, we sample each clip with 1616 frames as input and use the cross-entropy as fine-tuning objective. The epoch is set as 1010.

E.3 VLP Evaluation on EgoMCQ

In Tab. 10, we display EgoMCQ evaluation result of Frozen pretrained on different video-text datasets.

As shown, pretraining with EPIC-KITCHENS-100 dataset (1st-person view, 67.2K67.2\text{K} pairs) reach comparable performance with HowTo100M pretraining (3rd-person view, 136M136\text{M} noisy pairs), which demonstrates the major domain gaps. Besides, Frozen with CC3M+WebVid-2M pretraining reach significant improvement on the intra-video setting, but minor in inter-video. We speculate this due to CC3M+WebVid-2M dataset covering a wide range of appearance information but still less exploration in the fine-grained action e.g. human-object interaction.

E.4 Training Curves of EPIC-KITCHENS-100 video-text retrieval

In Fig. 12, we display training curves of EPIC-KITCHENS-100 video-text retrieval under different video-text pretraining, which also includes a baseline without video-text pre-training. We can found that: Variants with video-text pretraining have a faster rise in performance. Except for HowTo100M, which is similar to variant without video-text pretraining. Especially with EgoClip for egocentric pretraining, the VLP model achieves nearly convergent performance with only a small number of epochs (less than 2020). With EgoNCE as pretraining objective, this positive effect is further enhanced.

E.5 Results on test set of Natural Language Query

In Tab. 11, we found the similar conclusions in test set of Natural Language Query, pretraining with EgoClip and EgoNCE reach the optimum performance.

E.6 Results on test set of Moment Query

We further display the test set results of Moment Query in Tab. 12, pretraining with EgoClip and EgoNCE reach the best performance, 3.78%3.78\% on R@1@1 and 4.65%4.65\% on Avg mAP over the baseline.

E.7 Visualization

To intuitively understand the effect of egocentric pre-training, in Fig. 13, we compare the EPIC-KITCHENS-100 video-text retrieval results between our pre-training (EgoClip w/ EgoNCE) and CC3M+WebVid-2M pre-training, both fine-tuning with 1616 frames. The numbers after each narration represent the correlation scores between the query and the retrieval result, with 11 being the best.