Moments in Time Dataset: one million videos for event understanding
Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa Brown, Quanfu Fan, Dan Gutfruend, Carl Vondrick, Aude Oliva
Introduction
“The best things in life are not things, they are moments” of raining, walking, splashing, jumping, etc. Moments happening in the world unfold at time scales from a second to minutes, occur in different places, and involve people, animals, objects and natural phenomena. Of particular interest are moments of a few seconds as they represent an ecosystem of diverse visual and auditory dynamic events.
We introduce the Moments in Time Dataset, a collection of one million short videos each with a label corresponding to an event unfolding in 3 seconds.The website is http://moments.csail.mit.edu Temporal events of such length correspond to the average duration of human working memory which is a short-term memory-in-action buffer specialized in representing information that is changing over time. Additionally, three seconds is a temporal envelope which holds meaningful actions between people, objects and phenomena (e.g. wind blowing, objects falling on the floor, shaking hands, playing with a pet, etc).
Compound activities that occur at longer time scales can be represented by sequences of three second actions. For example, picking up an object and running could be interpreted as the compound actions "stealing", "saving" or "playing sports" depending on the context of the activity (e.g. agent and scene). Hypothetically, when describing such a "stealing" event, one can go into the details of the movement of each joint and limb of the persons involved. However, this is not how we naturally describe compound events. Instead, we use verbs such as "picking" and "running" which are the actions which typically occur in a time window of 1-3 seconds. The ability to automatically recognize these short actions is a core step for automatic video comprehension.
Modeling the spatial-temporal dynamics even for three second videos, poses a daunting challenge. For instance, videos with the action "opening" include people opening doors, gates, drawers, and curtains, animals and humans opening eyes, and even a flower opening its petals. In some cases the same set of frames in reverse can actually depict a different action ("closing") showing that the temporal aspect is crucial to video understanding. Humans can recognize a common transformation that occurs in space and time that allows for all of the mentioned scenarios to be assigned to the category "opening" even though visually they look very different from each other. The challenge is to develop models that recognize these transformations in a way that will allow them to discriminate between different actions, yet generalize to other agents and settings within the same class.
We present the Moments in Time dataset, one million videos each with one action label from 339 different classes, to enable models to richly understand actions and dynamics in videos. This is one of the largest human-annotated video datasets capturing visual and audible short events produced by humans, animals, objects or nature. The most commonly used verbs in the English language are chosen as the vocabulary covering a wide and diverse semantic space. This presents a uniquely challenging and diverse dataset for action recognition with significant intra-class variation. In Section 4 we include baseline results of several known models trained and tested on the dataset, addressing separately, and jointly, three modalities: spatial, temporal and auditory.
Related Work
Video Datasets: Over the years, the size of datasets for video understanding has grown steadily. KTH and Weizmann were early datasets for human action understanding. UCF101 and THUMOS are built from web videos and have become important benchmarks for video classification. Kinetics and YouTube-8M introduced a large number of event classes by leveraging public videos from YouTube. The micro-videos dataset uses social media videos to study an open-world vocabulary for video understanding. ActivityNet explores recognizing activities in video and AVA explores recognizing amd localizing fine-grained actions. The “something something” dataset and Charades used crowdsourced workers to collect video datasets while the VLOG dataset collects daily human activities with natural spatio-temporal context.
Video Classification: The availability of large-scale video datasets has enabled significant progress in video understanding and classification. In early work, Laptev and Lindeberg developed space-time interest point descriptors and Klaser et al. designed histogram features for video. Pioneering work by Wang et al. developed dense action trajectories by separating foreground motion from camera motion. Sadanand and Corso designed ActionBank as a high-level representation for video and action classification, and Pirsiavash and Ramanan leveraged grammar models for temporally segmenting actions from video. Advances in deep convolutional networks have enabled the development of a variety of large-scale video classification models. Various approaches of fusing RGB frames over the temporal dimension are explored on the Sport1M dataset . Two stream CNNs with one static image stream and one optical flow stream were proposed to fuse the information of object appearance and short-term motion . 3D convolutional networks use 3D kernels to extract features from a sequence of RGB frames. Temporal Segment Networks sample frames and optical flow on different time scales to extract information for activity recognition . A CNN+LSTM model, which uses a CNN to extract frame features and an LSTM to integrate features over time, is also used to recognize activities in videos . Recently, I3D networks use two stream CNNs with inflated 3D convolutions on both RGB and optical flow sequences to achieve state of the art results on the Kinetics dataset . More recent 3D networks incorporate Non-local modules in order to capture long-range dependencies while Temporal Relation Networks take a different approach by learning the relevant state transitions in sparsely sampled frames from different temporal segments.
Sound Classification: Environmental and ambient sound recognition is a rapidly growing area of research. Stowell et al. collected an early dataset and assembled a challenge for sound classification, Piczak collected a dataset of fifty sound categories and enough to train deep convolutional models, Salamon et al. released a dataset of urban sounds, and Gemmeke et al. use web videos for sound dataset collection. Recent work is now developing models for sound classification with deep neural networks. For example, Piczack pioneered early work for convolutional networks for sound classification, Aytar et al. transfer visual models into sound for auditory analysis, and Hershey et al. develop large-scale convolutional models for sound classification, and Arandjelović and Zisserman train sound and vision representations jointly. In Moments in Time dataset, many videos have both visual and auditory signals, enabling for multi-modal video recognition.
The Moments in Time Dataset
The goal of this project is to design a high-coverage, high-density, balanced dataset of hundreds of verbs depicting moments of a few seconds. High-quality datasets should have a broad coverage, high diversity and density of samples, and the ability to scale. The Moments in Time Dataset consists of over one million 3-second videos corresponding to 339 different verbs. Each verb is associated with over 1,000 videos resulting in a large balanced dataset for learning dynamic events from videos. Importantly, the dataset is designed to have, and grow towards, a very large set of both inter-class and intra-class variation that captures a dynamic event at different levels of abstraction (i.e. "opening" doors, curtains, eyes, mouths, even a flower opening its petals).
We began building our vocabulary by forming a list of the 4,500 most commonly used verbs from VerbNet (according to the word frequencies in the Corpus of Contemporary American English (COCA) ). We clustered these verbs using features from Propbank , FrameNet and OntoNotes which contain information on both the meaning and usage of each verb. This allows for clusters to be formed that extend beyond synonymous groups.
We formed our clusted by assigning a binary feature vector to each verb with an index for different features provided by FrameNet, PropBank and VerbNet. If a verb was associated with a given feature it’s value was set to 1 and 0 if not. These feature vectors were then used to perform k-means clustering to create the set of semantic verb clusters. For example, the actions "pour" and "flow" have similar meanings and can both be used to describe the movement of a liquid. However, "pour" has features associated with an agent (e.g. a person pouring coffee) whereas "flow" does not. These differences can lead to very different video clips when we take into account the full context of an event.
To order our list of verbs we iteratively selected the most common verb from the cluster with the highest cumulative frequency of use among its members and added it to our vocabulary. For example, a cluster associated with "grooming" contains the following verbs in order of most common to least common "washing, showering, bathing, soaping, grooming, shampooing, manicuring, moisturizing, and flossing". Verbs can belong to multiple clusters due to their different frames of use. For instance, "washing" also belongs to a group associated with cleaning, mopping, scrubbing, etc. Once a verb was chosen it was then removed from all of its member clusters. We repeated this process for all verbs in the set excluding verbs that were either ambiguous, not likely to be visual/audible in a 3-seconds video (e.g. "thinking" and "being") or too similar to a previously selected verb. We settled on a set of 339 frequently used and semantically diverse verbs that we used to build the proposed dataset with a large coverage and diversity of labels.
2 Collection and Annotation
To generate candidate videos for annotation, we search the Internet by parsing video metadata and crawling search engines to build a list of candidate videos for each class in our vocabulary using a variety of different sources Youtube, Flickr, Vine, Metacafe, Peeks, Vimeo, VideoBlocks, Bing, Giphy, The Weather Channel, and Getty-Images. We download each video and randomly cut a 3-second section which we group with the corresponding verb. These verb-video tuples are then sent to Amazon Mechanical Turk (AMT) for annotation. To ensure the highest level of diversity, we cut a single 3-second snippet from each video source. We used this approach instead of using a model to localize interesting video segments, such as Video2Gif , to reduce any model bias where a trained model will localize segments more similar to the data for which it was trained.
Each AMT worker is presented with a video-verb pair and asked to press a Yes or No key signifying if the action is happening in the scene. Positive responses from the first round are sent to subsequent rounds of annotation. Each HIT (a single worker assignment) contains 64 different 3-second videos that are related to a single verb and 10 ground truth videos that are used for control. In each HIT, the first 4 questions are used to train the workers on the task and require the correct answer to be selected before continuing. Only the results from HITs that earn a 90% or above on the control videos are included in the dataset. This binary-classification setup eases class selection for workers allowing for efficient annotation. We run each video in the training set through annotation at least 3 times and require a human consensus of at least 75% to be considered a positive label. For the validation and test set we increase the minimum number of rounds of annotation to 4 with a human consensus of at least 85%. We do not set the threshold at 100% to allow for videos with more difficult to recognize actions. Figure 2 shows an example of the annotation task.
3 Dataset Statistics
A motivation for this project was to gather a large balanced and diverse dataset for training models for video understanding. Since we pull our videos from over 10 different sources we are able to include a large breadth of diversity that would be challenging using a single source. In total, we have collected over 1,000,000 labelled videos for 339 Moment classes. The graph on the left of Figure 3 shows the full distribution across all classes where the average number of labeled videos per class is 1,757 with a median of 2,775.
To further aid in building a diverse dataset we do not restrict the active agent in our videos to humans. Many events such as "walking", "swimming", "jumping", and "carrying" are not specific to human agents. In addition, some classes may contain very few videos with human agents (e.g. "howling" or "flying"). True video understanding models should be able to recognize the event across agent classes. With this in mind we decided to build our dataset to be general across agents and present a new challenge to the field of video understanding. The middle graph in Figure 3 shows the distribution of the videos according to agent type (human, animal, object) for each class. On the far left (larger human proportion), we have classes such as "typing", "sketching", and "repairing", while on the far right (smaller human proportion) we have events such as "storming", "roaring", and "erupting".
Another feature of the Moments in Time dataset is that we include sound-dependant classes. We do not restrict our videos to events that can be seen, if there is a moment that can only be heard in the video (e.g. "clapping" in the background) then we still include it. This presents another challenge in that purely visual models will not be sufficient to completely solve the dataset. The right graph in Figure 3 shows the distribution of videos according to whether or not the event in the video can be seen.
4 Dataset Comparisons
In order to highlight the key points of our dataset, we compare the scale and object-scene coverage found in Moments in Time to other large-scale video datasets for action recognition. These include UCF101 , ActivityNet , Kinetics , Something-Something , AVA , and Charades . Figure 4 compares the total number of action labels used for training (left) and the average number of videos that belong to each class in the training set (middle). This increase in scale for action recognition is beneficial for training large generalizable systems for machine learning.
Additionally, we compared the coverage of objects and scenes that can be recognized within the videos. This type of comparison helps to showcase the visual diversity of our dataset. To accomplish this, we extract 3 frames from each video evenly spaced at 25%, 50%, and 75% of the video duration and run a 50 layer resnet trained on ImageNet and a 50 layer resnet trained on Places over each frame and average the prediction results for each video. We then compare the total number of objects and scenes recognized (top 1) by the networks in Figure 4 (right). The graph shows that 100% of the scene categories in Places and 99.9% of the object categories in ImageNet were recognized in our dataset. The closest dataset to ours in this comparison is Kinetics which has a recognized coverage of 99.5% of the scene categories in Places and 96.6% of the object categories in ImageNet. We should note that we are comparing the recognized categories from the top 1 prediction of each network. We have not annotated the scene locations and objects in each video of each dataset. However, a comparison of the visual features recognized by each network does still serve as an informative comparison of visual diversity.
Experiments
In this section we present the details of our experimental setup utilized to obtain the reported baseline results.
Data. We generate a training set of 802,264 videos with between 500 and 5,000 videos per class for 339 different classes and evaluate performance on a validation set of 33,900 videos with 100 videos for each class. We additionally withhold a test set of 67,800 videos consisting of 200 videos per class which will be used to evaluate submissions for a future action recognition challenge.
Preprocessing. We extract RGB frames from the videos at 25 fps and resize the RGB frames to a standard 340x256 pixels. In the interest of performance, we pre-compute optical flow on consecutive frames using an off-the-shelf implementation of TVL1 optical flow algorithm from the OpenCV toolbox . This formulation allows for discontinuities in the optical flow field and is thus more robust to noise. For fast computation, we discretize the values of optical flow fields into integers, clip the displacement with a maximum absolute value of 15 and scale the range to 0-255. The x and y displacement fields of every optical flow frame are then stored as two grayscale images to reduce storage. To correct for camera motion, we subtract the mean vector from each displacement field in the stack. For video frames, we use random cropping for data augmentation and subtract the ImageNet mean from images.
Evaluation metric. We use top-1 and top-5 classification accuracy as the scoring metrics. Top-1 accuracy indicates the percentage of testing videos for which the top confident predicted label is correct. Top-5 accuracy indicates the percentage of the testing videos for which the ground-truth label is among the top 5 ranked predicted labels. This is appropriate for video classification as videos may contain multiple actions (see Figure 5). For evaluation we randomly select 10 crops per frame and average the results.
2 Baselines for Video Classification
Here, we present several baselines for video classification on the Moments in Time dataset. We show results for three modalities (spatial, temporal, and auditory), as well as for recent video classification models such as Temporal Segment Networks and Temporal Relation Networks . We further explore combining models to improve recognition accuracy. The details of the baseline models grouped by different modalities are listed below.
Spatial modality. We experiment with a 50 layer resnet (Resnet50) trained on randomly selected RGB frames from each video with networks trained from scratch (ResNet50-scrach), initialized on Places (ResNet50-Places), and initialized on ImageNet (ResNet50-ImageNet). In testing, we average the prediction from 6 equi-distant frames.
Auditory modality. While many actions can be recognized visually, sound contains complementary or even mandatory information for recognition of particular classes, such as cheering or talking, as can be seen in Figure 3 (right). We use raw waveforms as the input modality and finetune a SoundNet network which was pretrained on 2 million unlabeled videos from Flickr with the output layer changed to predict moment classes (SoundNet).
Temporal modality. Following , we compute the optical flow between adjacent frames encoded in Cartesian coordinates as displacements by stacking together 5 consecutive frames to form a 10 channel image (the x and y displacement channels). We then modify the first convolutional layer of a BNInception model to accept 10 input channels (BNInception-Flow).
Spatial-Temporal modality. We also train three recent action recognition models: Temporal Segment Networks (TSN) , Temporal Relation Networks and Inflated 3D convolutional networks (I3D) . Temporal Segment Networks aim to efficiently capture the long-range temporal structure of videos using a sparse frame-sampling strategy. The TSN’s spatial stream TSN-Spatial is fused with an optical flow stream TSN-Flow via average consensus to form the two stream TSN TSN-2stream. The base model for each stream is a BNInception model with three time segments.
Temporal Relation Networks (TRN) explicitly learn temporal dependencies between video segments that best characterize a particular action. This “plug-and-play" module can simultaneously model several short and long range temporal dependencies to classify actions that unfold at multiple time scales. We trained a TRN with 8 multi-scale relations TRN-Multiscale on RGB frames using a Resnet50 base model. Note that we classify the TRN-Multiscale as spatiotemporal modality because in training it utilizes the temporal dependency of different frames.
Inflated 3D convolutional networks (I3D) inflate the convolutional and pooling kernels of a pretrained 2D network to a third dimension. The inflated 3D kernel is initialized from the 2D model by repeating the weights from the 2D kernel over the temporal dimension. This improves learning efficiency and performance as 3D models contain far more parameters than their 2D counterpart and a strong intialization greatly improves training. For our experiments we use a 3D Resnet50 inflated from our best 2D Resnet50 (spatial) from Table I. We train each model with 16 frames selected at 5 frames-per-second (fps) for each video.
Ensemble. To combine different modalities for action recognition, we form an ensemble using the top performing model of each modality (spatial: ResNet50-ImageNet, spatiotemporal: I3D and auditory: SoundNet). We concatenate the features from the final hidden layer of each modality and train a linear SVM to predict the moment categories (SVM). This ensemble learns how to fuse the features of the different modalities in order to make a more robust prediction based on spatial, temporal and auditory information giving us our highest recognition scores of 31.16% top-1 and 57.67% top-5.
3 Baseline Results
Table I shows the performance of the baseline models on the validation set. The best single model is I3D, with a Top-1 accuracy of 29.51% and a Top-5 accuracy of 56.06% while the Ensemble model (SVM) achieves a 57.67% Top-5 accuracy.
Figure 5 illustrates some of the high scoring predictions from the baseline models. These qualitative results suggest that the models can recognize moments well when the action is well-framed and close up. However, the model frequently misfires when the category is fine-grained or there is background clutter. Figure 6 shows examples where the ground truth category is not detected in the top-5 predictions due to either significant background clutter or difficulty in recognizing actions across agents.
We visualize the prediction given by the model by generating heatmaps for some video samples using Class Activation Mapping (CAM) in Figure 7. CAM highlights the most informative image regions relevant to the prediction. Here we use the top-1 prediction of the ResNet50-ImageNet model for each individual frame of the given video.
Categories that perform the best tend to have clear appearances and lower intra-class variation, for example bowling and surfing frequently happen in specific scene categories. The more difficult categories, such as covering, slipping, and plugging, tend to have wide spatiotemporal support as they can happen in most scenes and with most objects. Recognizing actions uncorrelated with scenes and objects seems to pose a challenge for video understanding.
Auditory models have qualitatively different performance per category versus visual models suggesting that sound provides a complementary signal to vision. However, the full ensemble model has per category performance that is fairly correlated with a single image (spatial) model. Given the relatively low performance on Moments in Time, this suggests that there is still room to capitalize on temporal and auditory dynamics to better recognize actions.
4 Cross Dataset Transfer
We additionally conducted a set of transfer experiments where we pretrain two models, one on Kinetics and one on Moments in Time, in order to evaluate which model generalizes better to other datasets. We use a Resnet50 I3D model for the experiments as this model gave us the best single stream performance on our dataset (Table I). We compare our results when transferring to UCF101 , HMDB51 and Something-Something . During training we randomly crop the average duration of each video in each dataset at 5 frames-per-second (fps). This allows for scalable transfer between datasets with different video lengths while keeping the frame-rates consistent with the dataset used for pretraining. For evaluation we apply the model using a sliding window at 5 fps with a 2 frame step-size and average the results.
Table II shows the results of the transfer task where the top-1 and top-5 scores are calculated by evaluating on the validation set of the dataset used to fine-tune the model. We can see from the results that pretraining on Moments in Time results in better performance when transferring to HMDB51 and pretraining on Kinetics gives stronger results when transferring to UCF101. This makes sense as UCF101 and Kinetics share many classes (e.g. "playing guitar", "skydiving", "walking dog", etc.) and both consist solely of YouTube videos, while HMDB51 is built from multiple sources and has a classes more similar to Moments in Time (e.g. "eating", "laughing", "pouring", etc.). The results on Something-Something show that pretraining on Moments in Time improves performance on a dataset designed for learning action concepts (e.g. picking something up). Additionally, the fact that each dataset has videos which are consistently longer than 3 seconds (significantly so for HMDB51) suggests that the 3 second length of the videos in the Moments in Time dataset does not hinder performance when applied to datasets with much longer videos.
Conclusion
We present the Moments in Time Dataset, a large-scale collection of three second videos covering a wide range of dynamic events involving different agents (people, animals, objects, and natural phenomena). We report results of several baseline models addressing separately, and jointly, three modalities: spatial, temporal and auditory. This dataset presents a difficult task for the field of computer vision as the labels correspond to different levels of abstraction (a verb like "falling" can apply to many different agents and scenarios and involve objects and scenes of different categories, see Figure 1). It will serve as a new challenge to develop models that can appropriately scale to the level of complexity and abstract reasoning needed to process each video.
Acknowledgements: This work was supported by the MIT-IBM Watson AI Lab, the Intelligence Advanced Research Projects Activity (IARPA) via Department of Interior / Interior Business Center (DOI/IBC) contract number D17PC00341 and the Toyota Research Institute / MIT CSAIL Joint Research Center.