Ego-Exo: Transferring Visual Representations from Third-person to First-person Videos
Yanghao Li, Tushar Nagarajan, Bo Xiong, Kristen Grauman
Introduction
(0cm,-16cm) In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
Egocentric video captured by wearable cameras offers a unique perspective into human behavior. It is the subject of a recent surge in research interest in first-person activity recognition , anticipation , and video summarization with many valuable future applications in augmented reality and robotics. Compared to third-person videos, egocentric videos show the world through a distinct viewpoint, encode characteristic egocentric motion patterns due to body and head movements, and have a unique focus on hands, objects, and faces, driven by the camera wearer’s attention and interaction with their surroundings.
However, these unique properties also present a fundamental challenge for video understanding. On the one hand, learning models purely from egocentric data are limited by dataset scale. Current egocentric video datasets are small (e.g., 90k clips in EPIC-Kitchens-100 vs. 650k in Kinetics-700 ) and lack diversity (e.g., videos only in kitchen scenes). On the other hand, a purely exocentric approach that uses more readily available third-person video—the status-quo for pre-training video models —ignores the unique properties of egocentric video and faces a major domain mismatch. Prior work has shown that this latter strategy, though popular, is insufficient: pre-training egocentric action recognition models with third-person data alone produces significantly worse results than pre-training with first-person data . In an attempt to bridge the domain gap, prior work explores traditional embedding learning or domain adaptation approaches , but they require paired egocentric and third-person videos that are either concurrently recorded or annotated for the same set of activities, which are difficult to collect and hence severely limit their scope.
Despite their differences, we hypothesize that the exocentric view of activity should in fact inform the egocentric view. First, humans are able to watch videos of other people performing activities and map actions into their own (egocentric) perspective; babies in part learn new skills in just this manner . Second, exocentric video is not devoid of person-centered cues. For example, a close-up instructional video captured from the third-person view may nonetheless highlight substantial hand-object interactions; or video captured with a hand-held phone may follow an event (e.g., a parade) as it unfolds with attentional cues related to a head-mounted camera.
Building on this premise, in this work we ask: “How can we best utilize current video datasets to pre-train egocentric video models?” Our key idea is to discover latent signals in third-person video that approximate egocentric-specific properties. To that end, we introduce a feature learning approach in which ego-video features are guided by both exo-video activity labels and (unlabeled) ego-video cues, to better align traditional third-person video pre-training with downstream egocentric video tasks.
Specifically, we introduce a series of ego-inspired tasks that require the video model to be predictive of manipulated objects, spatiotemporal hand-object interaction regions, and general egocentricity. Then we incorporate these tasks into training as knowledge-distillation losses to supplement an action classification pre-training objective on third-person video. See Fig. 1.
By design, our video models can continue to enjoy large amounts of labeled third-person training data, while simultaneously embedding egocentric signals into the learned features, making them a suitable drop-in replacement for traditional video encoders for egocentric video tasks. Finally, our approach does not assume any paired or activity-labeled egocentric videos during pre-training; the egocentric signals are directly inferred from third-person video.
Our experiments on three challenging egocentric video datasets show that our “Ego-Exo” framework learns strong egocentric feature representations from third-person video. On Charades-Ego , our model improves over models pre-trained on Kinetics—the standard pre-training and fine-tuning paradigm—by +3.26 mAP, and outperforms methods that specifically aim to bridge the domain gap between viewpoints. Finally, our pre-trained model achieves state-of-the-art results on EPIC-Kitchens-100 , the largest available first-person dataset.
Related Work
The unique viewpoint in egocentric video presents interesting research challenges including action recognition and anticipation , daily life summary generation , inferring body pose , and estimating gaze . Several egocentric video datasets have been created to support these challenges . Model architectures proposed for these tasks include multi-stream networks , recurrent networks , 3D conv nets and spatially grounded topological graph models .
These architectures vary significantly, but all use video encoders that are similarly pre-trained with third-person video datasets, despite being applied to egocentric video tasks. In contrast, we introduce key egocentric losses during exocentric video pre-training that bridge the domain gap when applied to downstream egocentric video tasks.
Joint first/third person video understanding
Several strategies have been proposed to address the domain gap between first and third person video. Prior work learns viewpoint-invariant representations using embedding learning methods, and applies them to action recognition , video summarization , image retrieval , person segmentation , and attention-driven gaze prediction . Image generation methods use generative adversarial frameworks to synthesize one viewpoint from the other. Viewpoint invariance has also been treated as a domain adaptation task in prior work, adapting third-person video models for overhead drone-footage . Other methods use egocentric video as a modality to supplement top-view footage to improve identification and tracking models .
The above methods require paired datasets that are either simultaneously recorded or that share the same labels for instances across viewpoints. In contrast, our method leverages only third-person video datasets, but is augmented with pseudo-labels (derived from first-person models) to learn egocentric video representations, thus circumventing both the need for first-person video during pre-training and the need for paired labeled data.
Knowledge distillation for video
In knowledge distillation (KD), one network is trained to reproduce the outputs of another . Distillation serves to compress models or incorporate privileged information from alternate tasks . In videos, KD can incorporate information from alternate modalities like audio , depth and flow , or a combination , and object-level information . In the context of self-supervised learning, prior work assigns weak image labels to video instances as supervision for video-level models . In contrast, we use inferred weak labels that are relevant to the egocentric domain, and we use them alongside third-person video labels, rather than in place of them during pre-training.
Egocentric cues in video understanding models
Egocentric video offers several unique cues that have been leveraged to improve video understanding models . These include attention mechanisms from gaze and motor attention , active object detection , and hands in contact . We are also interested in such important egocentric cues, but unlike prior work we do not train models on labeled egocentric video datasets to detect them. Instead, we embed labels predicted for these cues as auxiliary losses in third-person video models to steer feature learning towards egocentric relevant features. Unlike any of these prior methods, our goal is to leverage third-person video to pre-train first-person video models.
Ego-Exo Approach
Our goal is to learn egocentric video representations from third-person video datasets, by discovering and distilling important cues about hands, objects, and interactions (albeit in a different viewpoint) that are relevant to egocentric activity during pre-training. To do this, we automatically assign third-person video instances with various egocentric pseudo-labels that span simple similarity scores to complex spatiotemporal attention maps, and then we introduce auxiliary losses that force our features to be predictive of these pseudo-labels. On the one hand, our approach retains the benefits of large-scale third-person video and the original action classification task to guide general video feature learning. On the other hand, we steer feature learning towards better egocentric features using automatically generated egocentric labels, as opposed to collecting manually labeled instances.
In the following sections, we first describe the traditional video pretraining framework (Sec 3.1) and how we incorporate our auxiliary loss terms into it (Sec 3.2). Next we describe the three egocentric tasks we use, namely Ego-Score, Object-Score, and Interaction-Map (Sec 3.3). Finally, we present our full training and evaluation pipeline in Sec 3.4.
Video models benefit greatly from strong initializations. The standard procedure for training egocentric video models is thus to first pre-train models using large-scale third-person video datasets, and then fine-tune for a specific downstream task.
Once pre-trained, the backbone weights are retained, the head is replaced with a task-specific classifier, and the new network is trained with instances from a target egocentric dataset to predict egocentric video labels.
2 Ego-Exo pre-training
Third-person pre-training alone results in strong, general-purpose video features. However, it ignores important egocentric signals and introduces a domain gap that limits its utility for downstream egocentric tasks. We introduce auxiliary egocentric task losses to overcome this gap.
Specifically, along with datasets and , we assume access to off-the-shelf video models that address a set of egocentric video understanding tasks. For each task , the model takes as input a video (as either frames or clips) and generates predicted labels . We use these pre-trained models to associate egocentric pseudo-labels to the third-person video instances in . We stress that the videos in are not manually labeled for any task .
Note that these pseudo-labels vary in structure and semantics, ranging from scalar scores (e.g., to characterize how egocentric-like a third-person video is), categorical labels (e.g., to identify the manipulated objects in video) and spatiotemporal attention maps (e.g., to characterize hand-object interaction regions). Moreover, these labels are egocentric-specific, but they are automatically generated for third-person video instances. This diverse combination leads to robust feature learning for egocentric video, as our experiments will show. Once pre-trained, we can retain our enhanced backbone weights to fine-tune on an egocentric video task using data from .
3 Auxiliary egocentric tasks
A good egocentric video representation should be able to capture the underlying differences between first- and third-person videos, to discriminate between the two viewpoints. Based on this motivation, we design an Ego-Score task to characterize the egocentricity likelihood of the video.
For this, we train a binary ego-classifier on the Charades-Ego dataset , which has both egocentric and third-person videos of indoor activities involving object interactions. While the dataset offers paired instances showing the same activity from two views, our method does not use this pairing information or egocentric activity labels. It uses only the binary labels indicating if a sample is egocentric or exocentric. Please see Supp. for more training details and an ablation study about the pairing information.
We use this trained classifier to estimate the real-valued pseudo task-labels for each video in our pre-training framework described in Sec 3.2. We sample multiple clips from the same video and average their score to generate a video-level label. Formally, for a video with clips we generate scores:
where is a scalar temperature parameter, is the predicted logits from the ego-classifier , and is the class label.
Third-person videos display various egocentric cues, resulting in a broad distribution of values for Ego-Score, despite sharing the same viewpoint (details in Supp). This score is used as the soft target in the auxiliary task loss, which we predict using a video classification head :
Object-Score: Finding interactive objects.
In egocentric videos, interactions with objects are often central, as evident in popular egocentric video datasets . Motivated by this, we designate an Object-Score task for each video that encourages video representations to be predictive of manipulated objects.
Rather than require ground-truth object labels for third-person videos, we propose a simple solution that directly uses an off-the-shelf object recognition model trained on ImageNet to describe objects in the video. Formally, for a video with frames we average the predicted logits from across frames to generate the video-level Object-Score :
where is predicted logits for the class from the recognition model, and is the temperature parameter.
Similar to the Ego-Score, we introduce a knowledge-distillation loss during pre-training to make the video model predictive of the Object-Score using a module :
Interaction-Map: Discovering hand interaction regions.
The Object-Score attempts to describe the interactive objects. Here we explicitly focus on the spatiotemporal regions of interactions in videos. Prior work shows it is possible to recognize a camera wearer’s actions by attending to only a small region around the gaze point , as gaze often focuses on hand-object manipulation. Motivated by this, we introduce an Interaction-Map task to learn features that are predictive of these important spatiotemporal hand-object interaction regions in videos.
We adopt an off-the-shelf hand-object detector to detect hands and interacting objects. For each frame in a video, the hand detector predicts a set of bounding-box coordinates and associated confidence scores for detected hands. These bounding boxes are scaled to —the spatial dimensions of the video clip feature. We then generate a spatiotemporal hand-map where the score for each grid-cell at each time-step is calculated by its overlap with the detected hands at time :
where is the set of predicted bounding boxes that overlap with the -th grid-cell at time . The full hand-map label is formed by concatenating the per-frame labels over time. We generate the corresponding object-map analogously. See Fig 3 for an illustrative example.
We use the hand-map and object-map as the Interaction-Map pseudo-labels for the third-person videos during pre-training. We introduce two Interaction-Map prediction heads, and , to directly predict the Interaction-Map labels from clip features using a 3D convolution head:
where is -th frame from training clip , and and are the predicted hand-map and object-map scores.
We predict Interaction-Maps instead of directly predicting bounding boxes via standard detection networks for two reasons. First, detection architectures are not directly compatible with standard video backbones—they typically utilize specialized backbones and work well only with high resolution inputs. Second, predicting scores on a feature map is more aligned with our ultimate goal to improve the feature representation for egocentric video tasks, rather than train a precise detection model.
4 Ego-Exo training and evaluation
The three proposed ego-specific auxiliary tasks are combined together during the pre-training procedure to construct the final training loss:
Note that third-person video instances without hand-object interactions or salient interactive objects still contribute to the auxiliary loss terms, and are not ignored. Our distillation models approximate the responses of the pre-trained egocentric models as soft-targets instead of hard labels, offering valuable information about perceived egocentric cues, whether positive or negative for the actual label.
Training with our auxiliary losses results in features that are more suitable for downstream egocentric tasks, but it does not modify the network architecture itself. Consequently, after pre-training, our model can be directly used as a drop-in replacement for traditional video encoders, and it can be applied to various egocentric video tasks.
Experiments
Our experiments use the following datasets.
Kinetics-400 is a popular third-person video dataset containing 300k videos and spanning 400 human action classes. We use this dataset to pre-train all our models.
Charades-Ego has 68k instances spanning 157 activity classes. Each instance is a pair of videos corresponding to the same activity, recorded in the first and third-person perspective. Our method does not require this pairing, and succeeds even if no pairs exist (Supp.).
EPIC-Kitchens is an egocentric video dataset with videos of non-scripted daily activities in kitchens. It contains 55 hours of videos consisting of 39k action segments, annotated for interactions spanning 352 objects and 125 verbs. EPIC-Kitchens-100 extends this to 100 hours and 90k action segments, and is currently the largest annotated egocentric video dataset.
Due to its large scale and diverse coverage of actions, Kinetics has widely been adopted as the standard dataset for pre-training both first- and third-person video models . EPIC-Kitchens and Charades-Ego are two large and challenging egocentric video datasets that are the subject of recent benchmarks and challenges.
Evaluation metrics.
We pre-train all models on Kinetics, and fine-tune on Charades-Ego (first-person only) and EPIC-Kitchens for activity recognition. Following standard practice, we report mean average precision (mAP) for Charades-Ego and top-1 and top-5 accuracy for EPIC.
Implementation details.
We build our Ego-Exo framework on top of PySlowFast and use SlowFast video models as backbones with input frames and stride . We use a Slow-only ResNet50 architecture for all ablation experiments, and a SlowFast ResNet50/101 architecture for final results.
For our distillation heads and (Sec 3.3), we use a spatiotemporal pooling layer, followed by a linear classifier. We implement our Interaction-Map heads and as two 3D conv layers with kernel sizes and , and ReLU activation.
For our combined loss function (Eqn 7), we set the loss weights , and to respectively through cross-validation on the EPIC-Kitchens validation set (cross-validation on Charades-Ego suggested similar weights). The temperature parameter in Eqn 1 and Eqn 3 is set to 1. Training schedule and optimization details can be found in Supp.
1 Ego-Exo pre-training
We compare our pre-training strategy to these methods:
Scratch does not benefit from any pre-training. It is randomly initialized and directly fine-tuned on the target egocentric dataset.
Third-only is pre-trained for activity labels on Kinetics 400 . This represents the status-quo pre-training strategy for current video models.
First-only is pre-trained for verb/noun labels on EPIC-Kitchens-100 , the largest publicly available egocentric dataset.
Domain-adapt introduces a domain adaptation loss derived from gradients of a classifier trained to distinguish between first- and third-person video instances . This strategy has been used in recent work to learn domain invariant features for third-person vs. drone footage .
Joint-embed uses paired first- and third-person video data from Charades-Ego to learn viewpoint-invariant video models via standard triplet embedding losses . We first pre-train this model with Kinetics to ensure that the model benefits from large-scale pre-training.
Ego-Exo is pre-trained on Kinetics-400, but additionally incorporates the three auxiliary egocentric tasks (Sec 3.3) together with the original action classification loss, to learn egocentric-specific features during pre-training.
For this experiment, all models share the same backbone architecture (Slow-only, ResNet-50) and only the pre-training strategy is varied to ensure fair comparisons. Domain Adapt uses additional unlabeled egocentric data during pre-training, but from the same target dataset that the model will have access to during fine-tuning. Joint-embed uses paired egocentric and third-person data, an advantage that the other methods do not have, but offers insight into performance in this setting. Only First-only has access to ego-videos labeled for actions during pre-training.
Table 1 shows the validation performance of different pre-training strategies. Third-only benefits from strong initialization from large-scale video pre-training and greatly outperforms models trained from Scratch. First-only performs very poorly despite being pre-trained on the largest available egocentric dataset, indicating that increasing scale alone is not sufficient—the diversity of scenes and activities in third-person data plays a significant role in feature learning as well. Domain-adapt and Joint-embed both learn viewpoint invariant features using additional unlabeled egocentric data. However, the large domain gap and small scale of the paired dataset limit improvements over Third-only. Our Ego-Exo method achieves the best results on both Charades-Ego and EPIC-Kitchens. The consistent improvements (+1.54% mAP on Charades-Ego, and +1.64%/+1.97% on EPIC verbs/nouns) over Third-only demonstrate the effectiveness of our proposed auxiliary egocentric tasks during pre-training. This is a key result showing the impact of our idea.
Fig 4 shows a class-wise breakdown of performance on Charades-Ego compared to the Third-only baseline. Our pre-training strategy results in larger improvements on classes that involve active object manipulations.
2 Ablation studies
We next analyze the impact of each auxiliary egocentric task in our Ego-Exo framework. As shown in Table 2, adding the Ego-Score task improves performance on both EPIC-Kitchens tasks, while adding Object-Score and Interaction-Map consistently improves all results. This reveals that despite varying structure and semantics, these scores capture important underlying egocentric information to complement third-person pre-training, and further boost performance when used together.
Fig 5 shows instances from Kinetics based on our auxiliary pseudo-label scores combined with the weights in Eqn 7. Our score is highest for object-interaction heavy activities (e.g., top row: knitting, changing a tire), while it is low for videos of broader scene-level activities (e.g., bottom row: sporting events). Note that these videos are not in the egocentric viewpoint—they are largely third-person videos from static cameras, but are ego-like in that they prominently highlight important features of egocentric activity (e.g. hands, object interactions).
Adding auxiliary ego-tasks during fine-tuning.
Table 3 shows that while both the baseline and our method further improve by adding the auxiliary task during fine-tuning, our improvements (Ego-Exo + aux) are larger, especially on Charades-Ego. This is likely because our distillation heads benefit from training to detect hands and objects in large-scale third-person video prior to fine-tuning for the same task on downstream ego-datasets.
3 Comparison with state-of-the-art
Finally, we compare our method with state-of-the-art models, many of which use additional modalities (flow, audio) compared to our RGB-only models. We include three competitive variants of our model using SlowFast backbones: (1) Ego-Exo uses a ResNet50 backbone; (2) Ego-Exo* additionally incorporates our auxiliary distillation loss during fine-tuning.Same as Ego-Exo + aux in Table 3, but here with a SlowFast backbone (3) Ego-Exo*-R101 further uses a ResNet-101 backbone.
Table 4 compares our Ego-Exo method with existing methods. Our Ego-Exo and Ego-Exo* yield state of the art accuracy, improving performance over the strongest baseline by +2.11% and +3.26% mAP. We observe large performance gains over prior work, including ActorObserverNet and SSDA , which use joint-embedding or domain adaptation approaches to transfer third-person video features to the first-person domain. In addition, unlike the competing methods, our method does not require any egocentric data which is paired or shared category labels with third-person data during pre-training.
EPIC-Kitchens.
Table 6 compares our method to state-of-the-art models on the EPIC-Kitchens test set. Ego-Exo and Ego-Exo* consistently improve over SlowFast (which shares the same backbone architecture) for all categories on both seen and unseen test sets. Epic-Fusion uses additional optical flow and audio modalities together with RGB, yet Ego-Exo outperforms it on the top-1 metric for all categories. AVSlowFast also utilizes audio, but is outperformed by our model with the same backbone (Ego-Exo*-R101) on the S1 test set.Table 6 compares existing methods under a controlled setting: using a single model with RGB or RGB+audio as input, and only Kinetics/ImageNet for pre-training. Other reported results on the competition page may use extra modalities, larger pre-training datasets, or model ensemble schemes (e.g. the top ranking method ensembles 8 models). On EPIC-Kitchen-100 , as shown in Table 5, Ego-Exo consistently improves over the SlowFast baseline on all evaluation metrics, and Ego-Exo*-R101 outperforms all existing state-of-the-art methods.
Conclusion
We proposed a novel method to embed key egocentric signals into the traditional third-person video pre-training pipeline, so that models could benefit from both the scale and diversity of third-person video datasets, and create strong video representations for downstream egocentric understanding tasks. Our experiments show the viability of our approach as a drop-in replacement for the standard Kinetics-pretrained video model, achieving state-of-the-art results on egocentric action recognition on Charades-Ego and EPIC-Kitchens-100. Future work could explore alternate distillation tasks and instance-specific distillation losses to maximize the impact of third-person data for training egocentric video models.
Acknowledgments: UT Austin is supported in part by the NSF AI Institute and FB Cognitive Science Consortium.
References
Supplementary Material
This section contains supplementary material to support the main paper text. The contents include:
(§S1) Implementation details for three pre-trained egocentric models from Sec. 3.3.
(§S2) Implementation details for Kinetics pre-training presented in Sec. 3.1.
(§S3) Implementation details for fine-tuning on downstream egocentric datasets.
(§S4) Additional results on EPIC-Kitchens, Charades-Ego and Something-Something v2 datasets.
(§S5) Additional ablation studies, including ablations of the Interaction-map model, using as pre-trained models, appending features from , the impact of egocentric dataset scale on model performance, implicit pairing information in Ego-Scores and effect of using different backbones for Ego-Exo.
(§S6) Additional qualitative results, including distribution of Ego-Score over Kinetics, additional qualitative examples and class-wise breakdown of improvements for auxiliary tasks .
Supplementary video. A demonstration video shows animated version of video clips for the qualitative examples in §S6.
We provide additional implementation details for the task models used in Auxiliary egocentric tasks from Sec. 3.3.
We use a Slow-only model with a ResNet-50 backbone as the ego-classifier . Then, we train on the Charades-Ego dataset in which each instance is assigned with a binary label indicating if it is egocentric or exocentric. We take a Kinetics-pretrained model as initialization and train with 8 GPUs in 100 epochs. We adopt a cosine schedule for learning rate decaying with a base learning rate as 0.01 and the mini-batch size is 8 clips per GPU. To generate pseudo-labels for Kinetic videos, we sample clips for each video and generate our Ego-Score using Eqn 1.
We directly use an off-the-shelf object recognition model trained on ImageNet as . Specifically, we use a standard ResNet-152 network from Pytorch Hubhttps://pytorch.org/hub/pytorch_vision_resnet/. For each Kinetics video, we sample frames and generate Object-Score following Eqn 3.
We adopt a pre-trained hand-object detector to discover hand interaction regions. For the detected bounding-box for hands and interactive objects from Kinetics videos, we keep only high-scoring predictions and eliminate bounding boxes with confidence scores less than 0.5.
Note that all the three pre-trained egocentric models are either easy to access (off-the-shelf models and ) or easy to train (). Meanwhile, our auxiliary losses do not require the modification of the network, thus our model can be directly used as a drop-in replacement for downstream egocentric video tasks after pre-training.
S2 Details: Pre-training on Kinetics
We follow the training recipe in when training on Kinetics, and use the same strategy for both Slow-only and SlowFast backbones and different implemented methods.
All the models are trained from scratch for 200 epochs. We adopt a synchronized SGD optimizer and train with 64 GPUs (8 8-GPU machines). The mini-batch size is 8 clips per GPU. The baseline learning rate is set as 0.8 with a cosine schedule for learning rate decaying. We use a scale jitter range of pixels for input training clips. We use momentum of 0.9 and weight decay of .
S3 Details: Fine-tuning on Ego-datasets
During fine-tuning, we train methods using one machine with 8 GPUs on Charades-Ego. The initial base learning rate is set as 0.25 with a cosine schedule for learning rate decaying. We train the models for 60 epochs in total. Following common practice in , we uniformly sample 10 clips for inference. For each clip, we take 3 crops to cover the spatial dimensions. The final prediction scores are temporally max-pooled. All other settings are the same as those in Kinetics training.
EPIC-Kitchens.
We use a multi-task model to jointly train verb and noun classification with 8 GPUs on EPIC-Kitchens . The models are trained for 30 epochs with the base learning rate as 0.01. We use a step-wise decay of the learning rate by a factor of at epoch 20 and 25. During testing, we uniformly sample 10 clips from each video with 3 spatial crops per clip, and then average their predictions. All other settings are the same as those in Kinetics training.
For EPIC-Kitchens-100 , we take the same optimization strategy as EPIC-Kitchens , except training with 16 GPUs with a 0.02 base learning rate.
S4 Additional results
We report results of an Ensemble of four Ego-Exo models on EPIC-Kitchens in Table S1. Specifically, the Ensemble model includes Ego-Exo and Ego-Exo* with ResNet-50 and ResNet-101 backbones. As shown in Table S1, the Ensemble Ego-Exo further improves the performance in all categories, and consistently outperforms the Ensemble model of Epic-Fusion.
Table S2 shows the Ensemble results on EPIC-Kitchens-100. The Ensemble model of Ego-Exo outperforms the current best model on the leaderboardhttps://competitions.codalab.org/competitions/25923#results in all categories at the time of submission, especially on noun and action classes with +5% and +3.5% improvements on Overall Top-1 metric. Note that even without Ensemble, our single model already ranks the first on leaderboard and achieves better results than the best leaderboard model in most categories.
Charades-Ego.
We only train SlowFast and Ego-Exo methods on the egocentric videos from Charades-Ego in Table 4 of Sec 4. As Charades-Ego also provides third-person videos, we further jointly train first-person and third-person video classification during fine-tuning on Charades-Ego. In this setting, SlowFast and Ego-Exo achieve 25.06 and 28.32 mAP, respectively, with a ResNet-50 backbone. Hence our model further improves over that multi-task setting.
Something-Something V2 (SSv2).
SSv2 is a non-ego dataset with videos containing human object interactions. We further apply our Ego-Exo method on this dataset and find our Ego-Exo method improves over baseline Third-only (59.49% 60.41% in accuracy). Though our goal remains to address egocentric video, it does seem our method can even have impact beyond it and works for general interaction scenario.
S5 Additional Ablation studies
We ablate the Interaction-Map task by only using the Hand-Map and Object-Map scores in Eqn 6. As shown in Table S3, using Hand-Map or Object-Map alone consistently improves over the baseline (Third-only) while combining them (Interaction-Map) achieves the best results overall.
In section 3.3, we introduce several auxiliary egocentric tasks and distill information from them into the video model using auxiliary losses in our Ego-Exo pre-training framework. An alternative way to exploit these signals is to directly use these models () from auxiliary egocentric tasks as our pre-trained models, then fine-tune them on the egocentric datasets. Specifically, we take the ego-classifier and object recognition model as pre-trained models. We do not include hand-object detector here as the detection backbone is not compatible with the video backbone.
As shown in Table S4, though auxiliary task models capture specific egocentric properties, directly use them as pre-trained models is still insufficient. Our methods successfully embed the egocentric information from these auxiliary tasks into the video model through distillation losses, and still enjoy the strong representations learned from the large-scale third-person dataset.
Another alternative way to exploit these information from auxiliary tasks () is to directly use the extracted features on egocentric datasets using . Specifically, we extract the embeddings after the global pooling layer of the three auxiliary models and concatenate them with the ego models during fine-tuning. The results are shown in the ‘append‘ rows of Table S4. It indicates these baselines are less effective than the proposed distillation scheme in our Ego-Exo method.
Impact of the scale of egocentric datasets.
We study our model performance under varying scales of egocentric video supervision by using different percentages of videos in EPIC-Kitchens-100 . Fig S1 shows that our model consistently outperforms the baseline Third-only at all dataset scales, though both models perform worse with less egocentric videos during fine-tuning. When using only 10% of training data, our method improves over the Third-only by +2.68% and 1.99% on verb and noun tasks.
Implicit pairing information in Ego-Scores.
We train the Ego-Classifier for Ego-Scores on Charades-Ego . During training, we only use the binary label to indicate the instance is egocentric or not and don’t utilize any pairing information. However, Ego-Score might still contains some implicit pairing information. Here, we conduct a ablation study by only taking one view (either ego or non-ego instance) for each pair in Charades-Ego when training . Table S5 shows that methods without any implicit pairing information achieves similar performance. This demonstrates that the implicit pairing information is not critical for our Ego-Exo method.
Ego-Exo with different backbone networks.
Besides using Slow and SlowFast backbones in Sec 4, Table S6 further compares the results using I3D and TSM as the backbone structure. Ego-Exo achieves better results over the baseline Third-only on both Charades-Ego and Epic-Kitchen datasets. It indicates the versatility of our idea wrt the chosen backbone.
S6 Additional qualitative results
As mentioned in Sec. 3.3, though Kinetics videos are predominantly captured in the third-person perspective, the Ego-Score generated by the pretrained classifier is not trivially low for all video instances. Fig S2 plots the distribution of values this score takes. While a majority of instances have very low scores (not ego-like), a large number of instances prominently feature egocentric signals (image inset, right) and have higher scores.
We present a qualitative result corresponding to the ablation experiment in Table 2 in the main paper. Fig. S3 shows a venn diagram where each circle contains classes from Charades-Ego that a particular ablated model in Table 2 improves upon, over the baseline model. For example, the red circle is a model with only Ego-score (row 2, Table 2). The overlapping regions between two circles contain classes that are improved by both corresponding ablated models. Note that the three ablated models all contains some classes which are only improved by one particular ablated model, which suggests that three auxiliary tasks capture different egocentric properties.
Additional qualitative examples
In Fig. S4, we show additional examples of instances from Kinetics, sorted by the scores generated by our pre-trained egocentric models to supplement Fig. 5 in the main paper. The top two rows contain instances with high scores (more ego-like, more prominently features objects, and more hand-object interactions), while the bottom two rows feature instances with low scores. Note that the instances shown are frames from the corresponding video clips. Typically, video clips with more ego-like viewpoint and motions usually have higher Ego-Score. Please see the animated version of this figure in the supplementary video.