HierVL: Learning Hierarchical Video-Language Embeddings
Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, Kristen Grauman
Introduction
Understanding human activity in video is a fundamental vision problem with abundant applications in augmented reality, robotics, and information retrieval. The field has made exciting advances, from new models for recognition and self-supervised representations to major datasets . Nonetheless, activity understanding in video lags noticeably behind object understanding in images, where today’s AI models compete well with people.
One key reason for this discrepancy is the fact that whereas objects present themselves directly in the pixels—no subtext required—activity naturally has broad temporal context rooted in the human actor’s (latent) intentions. Not only does an activity stretch across video frames, but also its interpretation relies on the larger context of what the person is trying to accomplish. Thus, there is a natural hierarchy of information in video, starting with the short-term “what the person is literally doing right now” (e.g., reaching for the stove) and going all the way to the long-term “what the person aims to do” (e.g., cook dinner).
As a step towards capturing this hierarchy, we explore video-language representation learning. Video often has accompanying timestamped text, whether from spoken narrations in a how-to video , closed caption text and scripts , or deliberate text annotations . Existing video-language models learn a correspondence between the two modalities by matching short video segments with their text counterpart, typically with a learned embedding that produces a language-enriched video clip encoder. However, this standard approach risks capturing only the short-term actions. Granular comments such as “now I pour milk in the pan” or “he picked up a water hose” fail to capture the overall goal of the activity, like making a coffee or cleaning a car. As a result, at inference time their encodings for unseen videos can be myopic and miss sequential dependencies between observed events.
To tackle this problem, we introduce HierVL: a novel hierarchical video-language model that captures both short-term actions and long-term intents in video. Unlike standard video-language embeddings, our method aims to simultaneously capture the immediate observed actions as well as their contribution to the longer-term goal. To that end, given training video accompanied by timestamped clip-level text descriptions as well as global (video-level) text summaries, HierVL learns a video-text embedding for hierarchical temporal understanding using two layers of contrastive learning. The top (parent) layer encourages the aggregated video clips to be close to the overarching textual summary (e.g., he makes spaghetti dinner), while the bottom (child) layer trains individual clips to be similar to their respective descriptions (e.g., he turns on the cooker). See Fig. 1.
To our knowledge, ours is the first work to create a hierarchical video-language embedding. Our idea to blend abstract textual summaries with literal text descriptions is new. Furthermore, our model design addresses constituent technical challenges—namely, we circumvent the typical expense of long-term feature learning by using aggregation of short-term features, and we show how to jointly train with two levels of annotation in a way that staves off catastrophic forgetting of either layer.
This hierarchical training yields not only global video-level representations that capture long-term information (e.g., intent and temporal dependencies), but also clip-level video features that are more expressive than those traditionally learned via single-level schemes. This happens by means of our parent-child learning framework, which requires the aggregation of clip features within a video to match the long-term context captured by the summary.
We demonstrate our model by training with the narrations and summaries in the 3,670-hour egocentric video dataset Ego4D . We show that HierVL outperforms strong baselines and state-of-the-art methods for multiple video benchmarks, successfully transferring its pretrained representation for inference on Charades-Ego , EPIC-KITCHENS , and HowTo100M .Note that we do not need any text or summary annotations for these downstream datasets and tasks. We evaluate our representations on both hierarchy levels. In particular, at the time of submission, HierVL achieves state-of-the-art performance on Ego4D Long Term Anticipation (LTA), Charades-Ego Action Recognition, EPIC-KITCHENS-100 Multi-Instance Retrieval (zero-shot and fine-tuned settings), and HowTo100M Long Video Classification.
Related Work
Activity recognition and detection. Video understanding spans tasks like action recognition , action anticipation , procedure learning , and action localization . Various video datasets facilitate research in these directions, including Internet video collections like HowTo100M , YouCookII , and CrossTask , as well as freshly recorded datasets like CharadesEgo , EPIC-KITCHENS , and Ego4D . As a training resource, we use Ego4D , a large-scale diverse collection of in-the-wild wearable camera videos of daily-life activity around the world. The Ego4D videos have low-level text descriptions (“narrations”) of every action performed by the camera wearer, as well as video-level summaries, making them well-suited for our idea.
Long-form video representations. Longer videos introduce computational bottlenecks, making long-form video understanding challenging. There are several workarounds to make the task computationally feasible. Traditional methods include using pre-computed features that minimize backpropagation requirements or decreasing the frame-rate . Recent methods mitigate the computational requirements by creating a “feature-bank” or caching memory . Structured state space sequence models (S4) reduce the quadratic complexity of self-attention to linear, enabling efficient training of long-sequence tasks. Another promising approach is to aggregate fine-grained clip-level features into an overall video representation, as typically employed for video classification tasks. While all these methods are video-only, we propose a multi-modal long-form representation for both visual and textual modalities.
Joint video and language learning. The idea of projecting visual and language representations in the same embedding space is widely used for multi-modal understanding . Such joint representations enable several tasks, like language grounding in images , image captioning , and image retrieval , as well as text-to-video retrieval , video captioning , and video question answering . Several of these methods use contrastive learning (e.g., InfoNCE ) and match video clips (or images) with their narrations (or captions) in a self-supervised manner. The self-supervised model in uses both narrow and broad windows of visual and audio, and focuses on short-form video (e.g., Kinetics 5s clips). HERO uses a hierarchical loss between video clips (few seconds long) and their frames using only clip-level text, while enhances parent-level understanding for video-to-para retrieval and action recognition by concatenating text sentences to form (non-abstractive) paragraphs for hierarchical training.
All these methods only focus on localized narrations/captions. A single text sentence is matched to a clip that is typically a few seconds long. There are two reasons for choosing smaller temporal windows: a) the narrations typically span only a few seconds, and b) longer clips introduce computational overload that makes training difficult. In contrast, we devise a hierarchical approach to use both clip-level narrations spanning a few seconds and abstractive video-level summaries spanning several minutes. We show that clip feature aggregation makes learning computationally feasible, and that using such hierarchical text descriptions improve both clip-level and video-level tasks.
Technical Approach
We propose HierVL, a novel video-language model that captures both clip- and video-level relations. Fig. 2 overviews our method. Next, we describe the annotations (Sec. 3.1), formalize the embedding learning approach (Sec. 3.2), and discuss the feature aggregation strategy (Sec. 3.3). Finally, we describe the loss function (Sec. 3.4), training process (Sec. 3.5), and implementation details (Sec. 3.6).
Consider a hierarchically annotated video dataset, where is a long video, is a sequence of text narrations describing every atomic action in the video, and is a high-level text summary for the whole video. Notationally, is an ordered collection of short clips (each spanning a few seconds) and is an ordered collection of narrations . Note that there is no constraint on the temporal span of the video , but in our experiments they are typically in minutes. As an illustration, can be “he cleans the painting brush” or “he rubs the excess paint” whereas high-level summary will be “he was painting in a drawing room”. The clip contains a visual demonstration of the narration , whereas is an abstractive summary of the full video . The idea is for clip-level representations to capture fine-grained actions in a video, while video-level representations should capture the overall goal of the task.
We leverage the Ego4D dataset for training our model. Ego4D consists of 3,670 hours of wearable camera video of daily-life activity, as captured by 931 unique camera wearers around the world. Among the Ego4D annotations are text descriptions (“narrations”) of every action performed by the camera wearer, as well as video-level text summaries, which meet our requirements for and , respectively. The free-form narrations are written at timepoints selected by the annotators to capture every action performed. Specifically, annotators first watched a full 5-minute video and wrote a short 1-3 sentence summary for the overall activity and environment. Then annotators were asked to pretend they were describing everything occurring in the video to a friend on the phone who cannot see the video. The result is a temporally dense play-by-play description—13.2 sentences per minute on average, for a total of 3.85M sentences (see Appendix D in for details).
2 Hierarchical joint video and text embedding
In our hierarchical setup, we have short-term video segment and short-term text . We want to learn short-term representations and , which we refer to as the visual short-term features and the textual short-term features. At the long-term level, we have and as a collection of multiple and multiple , respectively. Simultaneously, we want to learn long-term representations and (referred to as long-term visual feature and long-term text feature, respectively). Finally, we have , the long-term summary feature, which is typically a few sentences long and hence is also encoded with .
The goal is to project into a common space such that semantically related features are close. Mathematically, for any suitably selected similarity metric and , we would like to fulfill a child-level matching constraint:
and , as well as parent-level matching constraints:
Overall, Eq. 1 implies corresponding short-term representations should have higher similarity than non-matching ones, Eq. 2 (and Eq. 3) implies a video (and narrations) should have a higher similarity with its summary than with other summaries. Note that since we project both short-term and long-term features into a common space, we are allowing features even at different hierarchical levels to come close in the embedding space if they are semantically similar.
3 Efficient long-term features via aggregation
Obtaining long-term features is challenging in both visual and text modalities. Directly computing a long-term visual feature requires more resources due to its large video size and often leads to inferior performance and memory overflows . Self-attention models are suitable architectures for capturing long-term dependencies, but they are challenging to apply to large collections of text sentences (e.g., long documents) due to quadratic dependence on the token sequence length in transformer models . Longformer mitigates this problem by multi-level global and local attentions.
Taking inspiration from these works in both visual and textual domains, we use aggregations of short-term features as long-term representations and . Following this strategy, we define the long-term visual representation as . Similarly, the long-term textual representation is defined as . We consider two aggregator functions . The first uses a self-attention transformer block in order to capture long-term dependencies over the entire video. We use positional encodings in order to provide the model with the ability to embed temporal order information in the video-level representation. We denote with HierVL-SA the variant of our model based on this self-attention aggregator. The second form of aggregation that we consider is simple average pooling (i.e., a parameter-free aggregator), which produces long-term features with equal contributions from all short-term features. This aggregator does not preserve order information. We name his version HierVL-Avg. We use the same aggregator in both modalities since and have the same dimensions (and, in fact, equal values for matching visual-text pairs in an ideal contrastive training).
4 Contrastive pretraining objective
As introduced previously, we learn the representations at two levels—child-level and parent-level . For child level representations, the pretraining objective is similar to prior work that relates short-term visual representations to short-term textual representations. In particular, we use a variant of EgoNCE , an action- and scene-aware variation of InfoNCE . EgoNCE groups similar actions as positives and temporally close distinct actions as hard negatives. In contrast, we omit the latter, since our hierarchical setup ought to bring together distinct actions with the same camera-wearer intent. Overall, the short-term pretraining objective is:
At the parent level, we use a similar pretraining objective between - and -. See Fig. 2 (bottom). As discussed in Sec. 3.3, we aggregate to obtain (and aggregate to get ). Since the short-term matching already contrasts and , we do not contrast and again at the parent-level. Overall, the long-term pretraining objective is where
and similarly for . For the parent-level feature, negatives for a summary text are both visual and textual representations chosen from outside the temporal span of .
5 Training strategy
So far, we discussed our approach for hierarchical video-language pretraining. To realize this setup, we employ a joint training approach. First, we train batches of short-term visual and textual pairs — thus training and . Subsequently, we train one batch of long-term features — thereby training and . Recall that and . Therefore, in this batch, we update the weights of as well as short-term and . The contrastive objective is detailed in Sec. 3.4.
The motivation behind training both levels of annotations together is to ensure the functions and optimize for both short-term and long-term features, i.e., both are influenced by the text summaries. Other alternatives are (a) using separate models for clip-level and video-level features, but that increases the parameters in the model and makes the training difficult (both in terms of convergence and GPU usage), and (b) training with only clip-level data and fine-tuning it for video-level (or vice-versa), but such strategies are known to lead to catastrophic forgetting .
Fig. 3 visualizes the learned features for 500 summary texts and their child narrations using our (left) and EgoVLP’s features (right). While summary features in EgoVLP are unrelated to the narrations, HierVL captures their natural hierarchy, as seen by the colors clustering together in the embedding space. This reshaping of the features reflects how our clip-level features convey context about the higher-level intent of the camera wearer.
6 Implementation Details
Network architecture. To learn the video feature extractor , we use a standard FrozenInTime video backbone, which is a slight deviation from TimeSformer and inspired from ViT . ViT-based vision transformers are frequently used as a feature extractor owing to their superior performance compared to other backbones. The video representation is learned from scratch; the output representation is the output of the final CLS token. We choose frames at 1 fps for short-term clips. Next, the text feature extractor is a DistillBERT architecture which achieves performance on-par with BERT but offers the benefit of being lighter.
Aggregator. Our HierVL-SA variant is implemented by means of a 6-layer self-attention block of the TimeSformer architecture and HierVL-Avg is averaging of features. In order to have a constant batch size, for both HierVL-SA and HierVL-Avg, we aggregate 16 short-term representations uniformly sampled from the entire video.
Training setup and parameters. We pretrain our architecture on 4 nodes, each with eight 32 GB NVIDIA V100 GPUs for 10 epochs for two days. We use AdamW optimizer with a learning rate of . We train one batch of video-level aggregation after every epoch of clip-level training. We use a batch size of 16 per GPU for short-term contrastive learning and 1 per GPU for long-term video-level contrastive learning. Recall that one video-level batch consists of 16 clips of the same video.
Experiments
We first pretrain our architecture with the setup and parameters discussed in Sec. 3.6 and report its results on multiple tasks aimed directly at gauging the quality of the learned video features (Sec. 4.1). Next, we show that our pretrained model improves the state of the art on a variety of downstream tasks covering both short- and long-term understanding (Sec. 4.2).
We use Ego4D for our contrastive pretraining. Ego4D has two-level hierarchical annotations—short-term step-by-step narrations and a long-term summary of the demonstration as observed by an annotator. We maintain the same training and validation split as in . Overall, there are M short-term narrations and K long-term summary annotations.
Pretraining evaluation tasks. We evaluate the quality of pretraining on three tasks defined on the Ego4D dataset: EgoMCQ (multiple-choice-question, introduced in EgoVLP ), as well as two new benchmarks that we propose — SummaryMCQ and ShuffleMCQ. In EgoMCQ, the model is given a narration prompt along with five candidate clips and must match the prompt with the correct video clip, with accuracy as the performance metric. Intra-video and Inter-video are two splits of the validation data where the candidate video clips are selected from the same or the other videos, respectively. SummaryMCQ mimics the video-language matching test of EgoMCQ but here the model is given a summary and five candidate long-term video options. The options are videos spanning the whole summary duration. While EgoMCQ validates clip-level performance, SummaryMCQ validates video-level performance. Finally, ShuffleMCQ is designed to evaluate temporal understanding: a summary text is given, and only the correct option maintains the temporal order among clips. The other four video options are generated by randomly reshuffling clips of the original video.
Comparison to EgoVLP. Our main comparison is to EgoVLP , since our model adopts the same architecture and uses its EgoNCE as the short-term loss in the objective. However, while our method leverages a hierarchical contrastive training that makes use of summary information, EgoVLP only focuses on short-term visual-textual correspondences. For SummaryMCQ, we use parameter-free averaging to compute the aggregate representation.
Table 1 shows the results.The first row corresponds to the numbers reported in EgoVLP and the second row corresponds to the numbers that we reproduced using the same codebase. We attribute the difference in performance to different hardware configurations. EgoVLP and both variants of our HierVL perform similarly on EgoMCQ, consistent with the fact this task requires short-term information only. In contrast, HierVL-SA obtains significantly better accuracy on the video-level (long-term) tasks, SummaryMCQ and ShuffleMCQ. Specifically, HierVL-SA outperforms EgoVLP by more than on SummaryMCQ. This highlights our model’s ability to capture long-term intent more effectively than the aggregated short-term features of EgoVLP. On ShuffleMCQ, both EgoVLP and HierVL-Avg are no better than chance (20%). This reflects how neither model captures the temporal order information that is essential to distinguish between the original summary and shuffled videos. Conversely, HierVL-SA exhibits stronger performance, producing a gain of 6.8% over these models (a relative gain of 34%). In short, our hierarchical learning shines for the long-term video tasks, successfully encoding the longer-term dependencies between events. We also observe that HierVL-SA outperforms EgoVLP with varying model sizes. Thus, further scaling models would not diminish the need for our architecture (see Supp).
Ablating design choices. The bottom portion of Table 1 includes several variants of our HierVL, in order to ablate the different design choices. Our proposed architecture has three distinct components: (a) a hierarchical model that operates at two levels (parent-level summaries and child-level narrations), (b) use of text summaries as a supervision, and (c) the joint training of these hierarchical annotations.
HierVL-w/o Joint is a variant used to investigate the effectiveness of joint training (component c). We start HierVL-w/o Joint with EgoVLP pretrained weights and train the whole network () using summaries only, i.e., without narrations. In this variant, the clip representations are indirectly supervised by means of the parent loss. We can see that while HierVL-w/o Joint achieves decent results on the two video-level tasks, its performance on EgoMCQ is much lower than that achieved by EgoVLP, which is its initialization. This suggests that summaries by themselves are not sufficient to supervise the learning of strong clip-level representations.
HierVL-w/o Hier uses (b, c) but not (a), i.e., we use summary supervision without a hierarchical model. We randomly assign the summary text annotation to one of the short-term segments. Importantly, this baseline uses the same amount of supervision as our proposed HierVL, yet it has overall lower performance (except for a marginal gain on EgoMCQ Inter-video). This highlights the effectiveness of our hierarchical training scheme.
HierVL-w/o Summ uses (a, c) but not (b), i.e., the supervision does not come from the summary text. Note, this represents the main idea from . The parent-level positives for contrastive learning are and . The objective of this ablation is to determine if high-level summaries are needed, or whether an aggregation of narrations can serve as a high-level representation. We observe that this variant is considerably less effective than HierVL-SA on the two video-level tasks of SummaryMCQ and ShuffleMCQ. This is an important result, as it suggests that the high-level intent expressed by the human annotator in the summary is effectively captured by HierVL-SA and this human supervision cannot be adequately replaced by an aggregation of short-term narrations.
Finally, HierVL-w/o SummNarr investigates the need for an additional text-only parent-level matching, as given in Eq. 3. This ablation checks the effect of only matching vs. matching both and . We see that imposing additional supervision between child and parent text features does increase the performance on all validation sets.
2 Downstream Evaluation
We evaluate the representation learned by HierVL on multiple downstream tasks.
In addition to Ego4D we use Charades-Ego , which consists of 7,860 videos recorded from both first and third person viewpoints, with 157 action classes; EPIC-Kitchens-100 , an egocentric video of 100 hours of unscripted activities in 45 home kitchens in 4 cities; and HowTo100M , a large-scale YouTube dataset covering K visual “how-to” tasks.
Downstream tasks.
Long-Term Anticipation (LTA). Ego4D’s LTA challenge requires the model to predict the next 20 actions given the current action (verb, noun). Metric is Edit Distance (ED) .
Action Recognition. Charades-Ego’s task requires predicting the action among 157 categories. Metric is mAP (mean average precision). We evaluate both the zero-shot and fine-tuned settings.
Multi-Instance Retrieval (MIR). EPIC-Kitchens-100’s MIR is a text-to-video and video-to-text retrieval task. Metrics are mAP and nDCG (normalized Discounted Cumulative Gain) for both VT and TV. We report their averages. Again, we evaluate in both zero-shot and fine-tuned settings.
Video Classification. To demonstrate the transfer ability of our pretraining, we perform linear probing on the most frequent 100 classes in HowTo100M. Metric is classification accuracy.
Throughout, we report relevant comparisons from the best existing methods in the literature, as well as the “w/o Hier” ablation, which uses the exact same summary data/supervision as HierVL, hence pinpointing the influence of our hierarchical training idea.
Ego4D LTA: Tab. 2 shows results on the test set of Ego4D LTA challenge. The models need to forecast the future actions, which is non-trivial even for humans. We improve the state of the art in both verb and noun predictions. Additionally, ours is the best performing method on the public leaderboard at the time of submission (in Tab. 2 we only compare with published works). HierVL-w/o Hier does not perform well despite also having access to the summaries, thus asserting the effectiveness of our hierarchical training. We use our learned representations and followed by a multi-headed decoder, as in the baseline . This result shows the effectiveness of both our learned feature aggregator (long-term) as well as short-term visual encoder .
Charades-Ego Action Recognition. Tab. 3 (top) shows the zero-shot results. EgoVLP reports overfitting when transferring from Ego4D to Charades-Ego and hence chooses another pretraining checkpoint. There is a significant gap in the performance between the two checkpoints. We report results on both—best performing pretraining checkpoint (denoted as PT ckpt) and the checkpoint chosen by EgoVLP (denoted as Task ckpt). Our model does not overfit when transferring to Charades-Ego; our performance on the corresponding checkpoints are and higher. In this downstream evaluation, only the short-term visual encoder (frozen) is required. Clearly, our hierarchical pretraining improves short-term features as well.
Tab. 3 (bottom) shows the fine-tuned results for the same task. Here, to compare against state-of-the-art methods, we fine-tune the model starting from our best pretrained checkpoint (having mAP for HierVL-SA). We outperform the current state-of-the-art EgoVLP . We fine-tune for this task, showing improvement in the short-term features. To our knowledge, ours is the best reported result for this dataset in the literature.
EPIC-Kitchens-100 Multi-Instance Retrieval. Tab. 4 (top) shows the zero-shot results. We observe a gain of mAP and increase between the best method and our HierVL-SA. Our HierVL-Avg is also slightly better than the state-of-the-art method. In this task, we use both the short-term encoders and (both frozen) and thus this experiment also validates our claim of improved short-term representations via hierarchical learning. Tab. 4 (bottom) shows our fine-tuning results for the same task. We fine-tune both and . We increase both metrics compared to the state-of-the-art.
HowTo100M Video Classification. Tab. 5 shows the results. In this linear probe setting, all of and are frozen and only one additional linear layer is trainable (trainable parameters K). We see that all of our learned representations are better than the baseline EgoVLP. Parameter-free averaging works well in video classification . Therefore, we add a special case of HierVL-SA where we retain the pretrained and replace SA with average. This additional experiment also shows the superiority of short-term features in HierVL-SA compared to HierVL-Avg.
Conclusion
We introduce a novel hierarchical video-language embedding. Whereas current embeddings are oblivious to the long-term activity intent, HierVL focuses on both short-term “what is the person doing now” and long-term “what the person aims to do”. Through extensive experiments, we show that this improves both short-term and long-term video understanding. Our model pushes the state-of-the-art on a variety of video challenges, including the overall best performance in the literature on Charades-Ego action recognition and Ego4D long-term anticipation.
Acknowledgements: We thank Ziad Al-Halah and Tushar Nagarajan for feedback on the manuscript. KG is paid as a research scientist at Meta. UT Austin is supported in part by the IFML NSF AI Institute and NSF-CCRI.