Action Genome: Actions as Composition of Spatio-temporal Scene Graphs
Jingwei Ji, Ranjay Krishna, Li Fei-Fei, Juan Carlos Niebles
Introduction
Video understanding tasks, such as action recognition, have, for the most part, treated actions and activities as monolithic events . Most models proposed have resorted to end-to-end predictions that produce a single label for a long sequence of a video and do not explicitly decompose events into a series of interactions between objects. On the other hand, image-based structured representations like scene graphs have cascaded improvements across multiple image tasks, including image captioning , image retrieval , visual question answering , relationship modeling and image generation . The scene graph representation, introduced in Visual Genome , provides a scaffold that allows vision models to tackle complex inference tasks by breaking scenes into its corresponding objects and their visual relationships. However, decompositions for temporal events has not been explored much , even though representing events with structured representations could lead to more accurate and grounded action understanding.
Meanwhile, in Cognitive Science and Neuroscience, it has been postulated that people segment events into consistent groups . Furthermore, people actively encode those ongoing activities in a hierarchical part structure — a phenomenon referred to as hierarchical bias hypothesis or event segmentation theory . Let’s consider the action of “sitting on a sofa”. The person initially starts off next to the sofa, moves in front of it, and finally sits atop it. Such decompositions can enable machines to predict future and past scene graphs with objects and relationships as an action occurs: we can predict that the person is about to sit on the sofa when we see them move in front of it. Similarly, such decomposition can also enable machines to learn from few examples: we can recognize the same action when we see a different person move towards a different chair. While that was a relatively simple decomposition, other events like “playing football”, with its multiple rules and actors, can involve multifaceted decompositions. So while such decompositions can provide the scaffolds to improve vision models, how is it possible to correctly create representative hierarchies for a wide variety of complex actions?
In this paper, we introduce Action Genome, a representation that decomposes actions into spatio-temporal scene graphs. Object detection faced a similar challenge of large variation within any object category. So, just as progress in 2D perception was catalyzed by taxonomies , partonomies , and ontologies , we aim to improve temporal understanding with Action Genome’s partonomy. Going back to the example of “person sitting on a sofa”, Action Genome breaks down such actions by annotating frames within that action with scene graphs. The graphs captures both the objects, person and sofa, and how their relationships evolve as the actions progress from person - next to - sofa to person - in front of - sofa to finally person - sitting on - sofa. Built upon Charades , Action Genome provides K object bounding boxes with M relationships across K video frames with action categories.
Most perspectives on action decomposition converge on the prototypical unit of action-object couplets . Action-object couplets refer to transitive actions performed on objects (e.g. “moving a chair” or “throwing a ball”) and intransitive self-actions (e.g. “moving towards the sofa”). Action Genome’s dynamic scene graph representations capture both such types of events and as such, represent the prototypical unit. With this representation, we enable the study for tasks such as spatio-temporal scene graph prediction — a task where we estimate the decomposition of action dynamics given a video. We can even improve existing tasks like action recognition and few-shot action detection by jointly studying how those actions change visual relationships between objects in scene graphs.
To demonstrate the utility of Action Genome’s event decomposition, we introduce a method that extends to a state-of-the-art action recognition model by incorporating spatio-temporal scene graphs as feature banks that can be used to both predict the action as well as the objects and relationships involved. First, we demonstrate that predicting scene graphs can benefit the popular task of action recognition by improving the state-of-the-art on the Charades dataset from to and to when using oracle scene graphs. Second, we show that the compositional understanding of actions induces better generalization by showcasing few-shot action recognition experiments, achieving mAP using as few as training examples. Third, we introduce the task of spatio-temporal scene graph prediction and benchmark existing scene graph models with new evaluation metrics designed specifically for videos. With a better understanding of the dynamics of human-object interactions via spatio-temporal scene graphs, we aim to inspire a new line of research in more decomposable and generalizable action understanding.
Related work
We derive inspiration from Cognitive Science, compare our representation with static scene graphs, and survey methods in action recognition and few-shot prediction.
Cognitive science. Early work in Cognitive Science provides evidence for the regularities with which people identify event boundaries . Remarkably, people consistently, both within and between subjects, carve out video streams into events, actions, and activities . Such findings hint that it is possible to predict when actions begin and end, and have inspired hundreds of Computer Vision datasets, models, and algorithms to study tasks like action recognition . Subsequent Cognitive and Neuroscience research, using the same paradigm, has also shown that event categories form partonomies . However, Computer Vision has done little work in explicitly representing the hierarchical structures of actions , even though understanding event partonomies can improve tasks like action recognition.
Action recognition in videos. Many research projects have tackled the task of action recognition. A major line of work has focused on developing powerful backbone models to extract useful representations from videos . Pre-trained on large-scale databases for action classification , these backbone models serve as cornerstones for downstream video tasks and action recognition on other datasets. To assist more complicated action understanding, another growing set of research explores structural information in videos including temporal ordering , object localization , and even implicit interactions between objects . In our work, we contrast against these methods by explicitly using a structured decomposition of actions into objects and relationships.
Table 1 lists some of the most popular datasets used for action recognition. One major trend of video datasets is providing considerably large amount of video clips with single action labels . Although these databases have driven the progress of video feature representation for many downstream tasks, the provided annotations treat actions as monolithic events, and do not study how objects and their relationships change during actions/activities. In the mean time, other databases have provided more varieties of annotations: AVA localizes the actors of actions, Charades contains multiple actions happening at the same time, EPIC-Kitchen localizes the interacted objects in ego-centric kitchen videos, DALY provides object bounding boxes and upper body poses for daily activities. Still, scene graph, as a comprehensive structural abstraction of images, has not yet been studied in any large-scale video database as a potential representation for action recognition. In this work, we present Action Genome, the first large-scale database to jointly boost research in scene graphs and action understanding. Compared to existing datasets, we provide orders of magnitude more object and relationship labels grounded in actions.
Scene graph prediction. Scene graphs are a formal representation for image information in a form of a graph, which is widely used in knowledge bases . Each scene graph encodes objects as nodes connected together by pairwise relationships as edges. Scene graphs have led to many state of the art models in image captioning , image retrieval , visual question answering , relationship modeling , and image generation . Given its versatile utility, the task of scene graph prediction has resulted in a series of publications that have explored reinforcement learning , structured prediction , utilizing object attributes , sequential prediction , few-shot prediction , and graph-based approaches. However, all of these approaches have restricted their application to static images and have not modelled visual concepts spatio-temporally.
Few-shot prediction. The few-shot literature is broadly divided into two main frameworks. The first strategy learns a classifier for a set of frequent categories and then uses them to learn the few-shot categories . For example, ZSL uses attributes of actions to enable few-shot . The second strategy learns invariances or decompositions that enable few-shot classification . OSS and TARN propose a measurement of similarity or distance measure between video pairs , CMN encodes uses a multi-saliency algorithm to encode videos , and ProtoGAN creates a prototype vector for each class . Our framework resembles the first strategy because we use the object and visual relationship representations learned using the frequent actions to identify few-shot actions.
Action Genome
Inspired from Cognitive Science, we decompose events into prototypical action-object units . Each action in Action Genome is representated as changes to objects and their pairwise interactions with the actor/person performing the action. We derive our representation as a temporally changing version of Visual Genome’s scene graphs . However, unlike Visual Genome, who’s goal was to densely represent a scene with objects and visual relationships, Action Genome’s goals is to decompose actions and as such, focuses on annotating only those segments of the video where the action occurs and only those objects that are involved in the action.
Annotation framework. Action Genome is built upon the Charades dataset , which contains action classes, of which are human-object activities. In Charades, there are multiple actions that might be occurring at the same time. We do not annotate every single frame in a video; it would be redundant as the changes between objects and relationships occur at longer time scales.
Figure 2 visualizes the pipeline of our annotation. We uniformly sample frames to annotate across the range of each action interval. With this action-oriented sampling strategy, we provide more labels where more actions occur. For instance, in the example, actions “sitting on a chair” and “drinking from a cup” occur together and therefore, result in more annotated frames, from each action. When annotating each sampled frame, the annotators hired were prompted with action labels and clips of the neighboring video frames for context. The annotators first draw bounding boxes around the objects involved in these actions, then choose the relationship labels from the label set. The clips are used to disambiguate between the objects that are actually involved in an action when multiple instances of a given category is present. For example, if multiple “cups” are present, the context disambiguates which “cup” to annotate for the action “drinking from a cup”.
Action Genome contains three different kinds of human-object relationships: attention, spatial and contact relationships (see Table 2). Attention relationships indicate if a person is looking at an object or not, and serve as indicators for which object the person is or will interacting with. Spatial relationships describe where objects are relative to one another. Contact relationships describe the different ways the person is contacting an object. A change in contact often indicates the occurrence of an actions: for example, changing from person - not contacting - book to person - holding - book may show an action of “picking up a book”.
It is worth noting that while Charades provides an injective mapping from each action to a verb, it is different from the relationship labels we provide. Charades’ verbs are clip-level labels, such as “awaken”, while we decompose them into frame-level human-object relationships, such as a sequence of person - lying on - bed, person - sitting on - bed and person - not contacting - bed.
Database statistics. Action Genome provides frame-level scene graph labels for the components of each action. Overall, we provide annotations for frames with a total of bounding boxes of object classes (excluding “person”), and instances of relationship classes. Figure 3 visualizes the log-distribution of object and relationship categories in the dataset. Like most concepts in vision, some objects (e.g. table and chair) and relationships (e.g. in front of and not looking at) occur frequently while others (e.g. twisting and doorknob) only occur a handful of times. However, even with such a distribution, almost all objects have at least k instances and every relationship as at least K instances.
Additionally, Figure 4 visualizes how frequently objects occur in which relationships. We see that most objects are pretty evenly involved in all three types of relationships. Unlike Visual Genome, where dataset bias provides a strong baseline for predicting relationships given the object categories, Action Genome does not suffer the same bias.
Method
We validate the utility of Action Genome’s action decomposition by studying the effect of simultaneously learning spatio-temporal scene graphs while learning to recognize actions. We propose a method, named Scene Graph Feature Banks (SGFB), to incorporate spatio-temporal scene graphs into action recognition. Our method is inspired by recent work in computer vision that uses the information “banks” . Information banks are feature representations that have been used to represent, for example, object categories that occur in the video , or even include where the objects are . Our model is most directly related to the recent long-term feature banks , which accumulates features of a long video as a fixed size representation for action recognition.
Overall, our SGFB model contains two components: the first component generates spatio-temporal scene graphs while the second component encodes the graphs to predict action labels. Given a video sequence , the aim of traditional multi-class action recognition is to assign multiple action labels to this video. Here, represents the video sequence made up of image frames . SGFB generates a spatio-temporal scene graph for every frame in the given video sequence. The scene graphs are encoded to formulate a spatio-temporal scene graph feature bank for the final task of action recognition. We describe the scene graph prediction and the scene graph feature bank components in more detail below. See Figure 5 for a high-level visualization of the model’s forward pass.
Previous research has proposed plenty of methods for predicting scene graphs on static images . We employ a state-of-the-art scene graph prediction method as the first step of our method. Given a video sequence , the scene graph predictor generates all the objects and connects each object with their relationships with the actor in each frame, i.e. . On each frame, the scene graph consists of a set of objects that a person is interacting with and a set of relationships . Here denotes the th relationship between the person with the object . Note that there can be multiple relationships between the person and each object, including attention, spatial, and contact relationships. Besides the graph labels, the scene graph predictor also outputs confidence scores for all predicted objects: and relationships: . We have experimented with various choices of and benchmark their performance on Action Genome in Section 5.3.
2 Scene graph feature banks
After obtaining the scene graph on each frame, we formulate a feature vector by aggregating the information across all the scene graphs into a feature bank. Let’s assume there are classes of objects and classes of relationships. In Action Genome, and . We first construct a confidence matrix with dimension , where each entry corresponds to an object-relationship category pair. We compure every entry of this matrix using the scores output by the scene graph predictor . . Intuitively, is a high value when is confident that there is an object in the current frame and its relationship with the actor is . We flatten the confidence matrix as the feature vector for each image.
Formally, is a sequence of scene graph features extracted from frames . We aggregate the features across the frames using methods similar to long-term feature banks , i.e. are combined with 3D CNN features extracted from a short-term clip using feature bank operators (FBO), which can be instantiated as mean/max pooling or non-local blocks . The 3D CNN embeds short-term information into while provides contextual information, critical in modeling the dynamics of complex actions with long time span. The final aggregated feature is then used to predict action labels for the video.
Experiments
Action Genome’s representation enables us to study few-shot action recognition by decomposing actions into temporally changing visual relationships between objects. It also allows us to benchmark whether understanding the decomposition helps improve performance in action recognition or scene graph prediction individually. To study these benefits afforded by Action Genome, we design three experiments: action recognition, few-shot action recognition, and finally, spatio-temporal scene graph prediction.
We expect that grounding the components that compose an action — the objects and their relationships — will improve our ability to predict which actions are occurring in a video sequence. So, we evaluate the utility of Action Genome’s scene graphs on the task of action recognition.
Problem formulation. We specifically study multi-class action recognition on the Charades dataset . The Charades dataset contains crowdsourced videos with a length of seconds on average. At any frame, a person can perform multiple actions out of a nomenclature of classes. The multi-classification task provides a video sequence as input and expects multiple action labels as output. We train our SGFB model to predict Charades action labels during test time and during training, provide SGFB with spatio-temporal scene graphs as additional supervision.
Baselines. Previous work has proposed methods for multi-class action recognition and benchmarked on Charades. Recent state-of-the-art methods include applying I3D and non-local blocks as video feature extractors (I3D+NL), spatio-temporal region graphs (STRG) , Timeception convolutional layers (Timeception) , SlowFast networks (SlowFast) , and long-term feature banks (LFB) . All the baseline methods are pre-trained on Kinetics-400 and the input modality is RGB.
Implementation details. SGFB first predicts a scene graph on each frame, then constructs a spatio-temporal scene graph feature bank for action recognition. We use Faster R-CNN with ResNet-101 as the backbone for region proposals and object detection. We leverage RelDN to predict the visual relationships. Scene graph prediction is trained on Action Genome, where we follow the same train/val splits of videos as the Charades dataset. Action recognition uses the same video feature extractor, hyper-parameters, and solver schedulers as long-term feature banks (LFB) for a fair comparison.
Results. We report performance of all models using mean average precision (mAP) on Charades validation set in Table 3. By replacing the feature banks with spatio-temporal scene graph features, we outperform the state-of-the-art LFB by mAP. Our features are smaller in size ( in SGFB versus in LFB) but concisely capture the more information for recognizing actions.
We also find that improving object detectors designed for videos can further improve action recognition results. To quantitatively demonstrate the potential of better scene graphs on action recognition, we designed an SGFB Oracle setup. The SGFB Oracle assumes that a perfect scene graph prediction method is available. The spatio-temporal scene graph feature bank therefore, directly encodes a feature vector from ground truth objects and visual relationships for the annotated frames. Feeding such feature banks into the SGFB model, we observe a significant improvement on action recognition: increase on mAP. Such a boost in performance shows the potential of Action Genome and compositional action understanding when video-based scene graph models are utilized to improve scene graph prediction. It is important to note that the performance by SGFB Oracle is not an upper bound on performance since we only utilize ground truth scene graphs for the few frames that have ground truth annotations.
2 Few-shot action recognition
Intuitively, predicting actions should be easier from a symbolic embedding of scene graphs than from pixels. When trained with very few examples, compositional action understanding with additional knowledge of scene graphs should outperform methods that treat actions as monolithic concept. We showcase the capability and potential of spatio-temporal scene graphs to generalize to rare actions.
Problem formulation. In our few-shot action recognition experiments on Charades, we split the action classes into a base set of classes and a novel set of classes. We first train a backbone feature extractor (R101-I3D-NL) on all video examples of the base classes, which is shared by the baseline LFB, our SGFB, and SGFB oracle. Next, we train each model with only examples from each novel class, where , for 50 epochs. Finally, we evaluate the trained models on all examples of novel classes in the Charades validation set.
Results. We report few-shot experiment performance in Table 4. SGFB achieves better performance than LFB on all -shot experiments. Furthermore, if with ground truth scene graphs, SGFB Oracle shows a -shot mAP improvement. We visualize the comparison between SGFB and LFB in Figure 6. With the knowledge of spatio-temporal scene graphs, SGFB better captures action concepts involving the dynamics of objects and relationships.
3 Spatio-temporal scene graph prediction
Progress in image-based scene graph prediction has cascaded to improvements across multiple Computer Vision tasks, including image captioning , image retrieval , visual question answering , relationship modeling and image generation . In order to promote similar progress in video-based tasks, we introduce the complementary of spatio-temporal scene graph prediction. Unlike image-based scene graph prediction, which only has a single image as input, this task expects a video as input and therefore, can utilize temporal information from neighboring frames to strength its predictions. In this section, we define the task, its evaluation metrics and report benchmarked results from numerous recently proposed image-based scene graph models applied to this new task.
Problem formulation. The task expects as input a video sequence where represents image frames from the video. The task requires the model to generate a spatio-temporal scene graph per frame. is represented as objects with category labels and bounding box locations. represents the relationships between objects and .
Evaluation metrics. We borrow the three standard evaluation modes for image-based scene graph prediction : (i) scene graph detection (SGDET) which expects input images and predicts bounding box locations, object categories, and predicate labels, (ii) scene graph classification (SGCLS) which expects ground truth boxes and predicts object categories and predicate labels, and (iii) predicate classification (PREDCLS), which expects ground truth bounding boxes and object categories to predict predicate labels. We refer the reader to the paper that introduced these tasks for more details . We adapt these metrics for video, where the per-frame measurements are first averaged in each video as the measurement of the video, then we average video results as the final result for the test set.
Baselines. We benchmark the following recently proposed image-based scene graph models for the task of spatio-temporal scene graph prediction: VRD’s visual module (VRD) , iterative message passing (IMP) , multi-level scene description network (MSDN) , graph convolution R-CNN (Graph R-CNN) , neural motif’s frequency prior (Freq-prior) , and relationship detection network (RelDN) .
Results. To our surprise, we find that IMP, which was one of the earliest scene graph prediction models actually outperforms numerous more recently proposed methods. The most recently proposed scene graph model, RelDN marginally outperforms IMP, suggesting that modeling similarlities between object and relationship classes improve performance in our task as well. The small gap in performance between the task of PredCls and SGCls suggests that these models suffer from not being able to accurately detect the objects in the video frames. Improving object detectors designed specifically for videos could improve performance. The models were trained only using Action Genome’s data and not finetuned on Visual Genome , which contains image-based scene graphs, or on ActivityNet Captions , which contains dense captioning of actions in videos with natural language paragraphs. We expect that finetuning models with such datasets would result in further improvements.
Future work
With the rich hierarchy of events, Action Genome not only enables research on spatio-temporal scene graph prediction and compositional action recognition, but also promises various research directions. We hope future work will develop methods for the following:
Spatio-temporal action localization. The majority of spatio-temporal action localization methods focus on localizing the person performing the action but ignore the objects, which are also involved in the action, that the person interacts with. Action Genome can enable research on localization of both actors and objects, formulating a more comprehensive grounded action localization task. Furthermore, other variants of this task can also be explored; for example, a weakly-supervised localization task where a model is trained with only action labels but tasked with localizing the actors and objects.
Explainable action models. Explainable visual models is an emerging field of research. Amongst numerous techniques, saliency prediction has emerged as a key mechanism to interpret machine learning models . Action Genome provides frame-level labels of attention in the form of objects that a the person performing the action is either looking at or interacting with. These labels can be used to further train explainable models.
Video generation from spatio-temporal scene graphs. Recent studies have explored image generation from scene graphs . Similarly, with a structured video representation, Action Genome enables research on video generation from spatio-temporal scene graphs.
Conclusion
We introduce Action Genome, a representation that decomposes actions into spatio-temporal scene graphs. Scene graphs explain how objects and their relationships change as an action occurs. We demonstrated the utility of Action Genome by collecting a large dataset of spatio-temporal scene graphs and used it to improve state of the art results for action recognition as well as few-shot action recognition. Finally, we benchmarked results for the new task of spatio-temporal scene graph prediction. We hope that Action Genome will inspire a new line of research in more decomposable and generalizable video understanding.
Acknowledgement. This work has been supported by Panasonic. This article solely reflects the opinions and conclusions of its authors and not Panasonic or any entity associated with Panasonic.