Hierarchical Object-oriented Spatio-Temporal Reasoning for Video Question Answering
Long Hoang Dang, Thao Minh Le, Vuong Le, Truyen Tran
Introduction
Much of the recent impressive progress of AI can be attributed to the availability of suitable large-scale testbeds. A powerful testbed – largely under-explored – is Video Question Answering (Video QA). This task demands a wide range of cognitive capabilities including learning and reasoning about objects and dynamic relations in space-time, in both visual and linguistic domains. A major challenge of reasoning over video is extracting question-relevant high-level facts from low-level moving pixels over an extended period of time. These facts include objects, their motion profiles, actions, interactions, events, and consequences distributed in space-time. Another challenge is to learn the long-term temporal relation of visual objects conditioning on the guidance clues from the question – effectively bridging the semantic gulf between the two domains. Finally, learning to reason from relational data is an open problem on its own, as it pushes the boundary of learning from simple one-step classification to dynamically construct question-specific computational graphs that realize the iterative reasoning process.
A highly plausible path to tackle these challenges is via object-centric learning since objects are fundamental to cognition . Objects pave the way towards more human-like reasoning capability and symbolic computing . Unlike objects in static images, objects in video have unique evolving lives throughout space-time. As object lives throughout the video, it changes its appearance and position, and interacts with other objects at arbitrary time. When observed in the videos, all these behaviors play out on top of a background of rich context of the video scene. Furthermore, in the question answering setting, these object-oriented information must be considered from the specific view point set by the linguistic query. With these principles, we pinpoint the key to Video QA to be the effective high-level relational reasoning of spatio-temporal objects under the video context and the perspective provided by the query (see Fig. 1). This is challenging for the complexity of the video spatio-temporal structure and the cross-domain compatibility gap between linguistic query and visual objects.
Toward such challenge we design a general-purpose neural unit called Object-oriented Spatio-Temporal Reasoning (OSTR) that operates on a set of video object sequences, a contextual video feature, and an external linguistic query. OSTR models object lifelong interactions and returns a summary representation in the form of a singular set of objects. The specialties of OSTR are in the partitioning of intra-object temporal aggregation and inter-object spatial interaction that leads to the efficiency of the reasoning process. Being flexible and generic, OSTR units are suitable building blocks for constructing powerful reasoning models.
For Video QA problem, we use OSTR units to build up Hierarchical Object-oriented Spatio-Temporal Reasoning (HOSTR) model. The network consists of OSTR units arranged in layers corresponding to the levels of video temporal structure. At each level, HOSTR finds local object interactions and summarizes them toward a higher-level, longer-term representation with the guidance of the linguistic query.
HOSTR stands out with its authentic and explicit modeling of video objects leading to the effective and interpretable reasoning process. The hierarchical architecture also allows the model to efficiently scale to a wider range of video formats and lengths. These advantages are demonstrated in a comprehensive set of experiments on multiple major Video QA tasks and datasets.
In summary, this paper makes three major contributions: (1) A semantic-rich object-oriented representation of videos that paves the way for spatio-temporal reasoning (2) A general-purpose neural reasoning unit with dynamic object interactions per context and query; and (3) A hierarchical network that produces reliable and interpretable video question answering.
Related Work
Video QA has been developed on top of traditional video analysis schemes such as recurrent networks of frame features or 3D convolutional operators . Video representations are then fused with or gated by the linguistic query through co-attention , hierarchical attention , and memory networks . More recent works advance the field by exploiting hierarchical video structure or separate reasoning out of representation learning . A share feature between these works is considering the whole video frames or segments as the unit component of reasoning. In contrast, our work make a step further by using detail objects from the video as primitive constructs for reasoning.
Object-centric Video Representation inherits the modern capability of object detection on images and continuous tracking through temporal consistency . Tracked objects form tubelets whose representation contributes to breakthroughs in action detection and event segmentation . For tasks that require abstract reasoning, the connection between objects beyond temporal object permanence can be established through relation networks . The concurrence of objects’ 2D spatial- and 1D temporal- relations naturally forms a 3D spatio-temporal graphs . This graph can be represented as either a single flattened one where all parts connect together , or separated spatial- and temporal-graphs . They can also be approximated as a dynamic graph where objects live through the temporal axis of the video while their properties and connection evolve .
Object-based Video QA is still in infancy. The works in and extract object features and feed them to generic relational engines without prior structure of reasoning through space-time. At the other extreme, detected objects are used to scaffold the symbolic reasoning computation graph which is explicit but limited in flexibility and cannot recover from object extraction errors. Our work is a major step toward the object-centric reasoning with the balance between explicitness and flexibility. Here video objects serve as active agents which build up and adjust their interactions dynamically in the spatio-temporal space as instructed by the linguistic query.
Method
Given a video and linguistic question , our goal is to learn a mapping function that returns a correct answer from an answer set as follows:
In this paper, a video is abstracted as a collection of object sequences tracked in space and time. Function is designed to have the form of a hierarchical object-oriented network that takes the object sequences, modulates them with the overall video context, dynamically infers object interactions as instructed by the question so that key information regarding arises from the mix. This object-oriented representation is presented next in Sec. 3.2, followed by Sec. 3.3 describing the key computation unit and Sec. 3.4 the resulting model.
2 Data Representation
In practice, we use Faster R-CNN with ROI-pooling to extract positional and appearance features; and DeepSort for multi-object tracking. We assume that the objects live from the beginning to the end of the video. Occluded or missed objects have their features marked with special null values which will be specially dealt with by the model.
2.2 Joint Encoding of “What” and “Where”
Since the appearance of an object may remain relatively stable over time while constantly changes, we must find joint positional-appearance features of objects to make them discriminative in both space and time. Specifically, we propose the following multiplicative gating mechanism to construct such features:
where serves as a position gate to (softly) turn on/off the localized appearance features . We choose and , where and are network weights with height of .
Along with the object sequences, we also maintain global features of video frames which hold the information of the background scene and possible missed objects. Specifically, for each frame , we form the the global features as the combination of the frame’s appearance features (pretrained ResNet vectors) and motion feature (pretrained ResNeXt-101) extracted from such frame.
2.3 Linguistic Representation
3 Object-oriented Spatio-Temporal Reasoning (OSTR)
With the videos represented as object sequences, we need to design a scalable reasoning framework that can work natively on the structures. Such a framework must be modular so it is flexible to different input formats and sizes. Toward this goal, we design a generic reasoning unit called Object-oriented Spatio-Temporal Reasoning (OSTR) that operates on this object-oriented structure and supports layering and parallelism.
Across the space-time domains, real-world objects have distinctive properties (appearance, position, etc.) and behaviors (motion, deformation, etc.) throughout their lives. Meanwhile, different objects living in the same period can interact with each other. The OSTR closely reflects this nature by containing the two main components: (1) Intra-object temporal attention and (2) Inter-object interaction (see Fig. 2).
The goal of the temporal attention module is to produce a query-specific summary of each object sequence into a single vector . The attention weights are driven by the query to reflect the fact that the relevance to the query varies across the sequence. In details, the summarized vector is calculated by
where is the Hadamard product, are learnable weights, is a binary mask vector to handle the null values caused by missed detections as mentioned in Sec. 3.2.
When the sequential structure is particularly strong, we can optionally employ a BiLSTM to model the sequence. We can then either utilize the last state of the forward LSTM and the first state of the backward LSTM, or place attention across the hidden states instead of object feature embeddings.
3.2 Inter-object Interaction
Here is the relevance of object w.r.t. the query . The norm operator is implemented as a softmax function over objects in our implementation.
where is a nonlinear activation (ELU in our implementation). After a fixed number of GCN iterations, the hidden states of the final layer are gathered as .
To recover the underlying background scene information and compensate for possible undetected objects, we augment the object representations with the global context :
These vectors line up to form the final output of the OSTR unit as a set of objects
4 Hierarchical Object-oriented Spatio-Temporal Reasoning (HOSTR)
Even though the partitioning of temporal and spatial interaction in OSTR brings the benefits of efficiency and modularity, such separated treatment can cause the loss of spatio-temporal information, especially with long sequences. This limitation prevents us from using OSTR directly on the full video object sequences. To allow temporal and spatial reasonings to cooperate along the way, we break down a long video into multiple short (overlapping) clips and impose an hierarchical structure on top. With such division, the two types of interactions can be combined and interleaved across clips and allow full spatio-temrporal reasoning.
Based on this motive, we design a novel hierarchical structure called Hierarchical Object-oriented Spatio-Temporal Reasoning (HOSTR) that follows the video multi-level structure and utilizes the OSTR units as building blocks. Our architecture shares the design philosophy of hierarchical reasoning structures with HCRN as well as the other general neural building blocks such as ResNet and InceptionNet. Thanks to the genericity of OSTR, we can build a hierachy of arbitrary depth. For concreteness, we present here a two-layer HOSTR corresponding to the video structure: clip-level and video-level (see Fig. 3).
In particular, we first split all object sequences constructed in Sec. 3.2 into equal-sized chunks corresponding to the video clips, each of frames. As the result, each chunk includes the object subsequences , where are the object features extracted in Eq. 2, and is the starting time of clip .
Similarly, we divide the sequence of global frame features into parts corresponding to the video clips. The global context for clip is derived from each part by an identical operation with the temporal attention for objects in Eqs. 3,4: .
Clip-level OSTR units work on each of these subsequences , context and query , and generate the clip-level representation of the chunk :
Outputs of the clip-level OSTRs are different sets of objects whose identities were maintained. Therefore, we can easily chain these objects of the same identity from different clips together to form the video-level sequence of objects .
At the video level, we have a single OSTR unit that takes in the object sequence , query and video-level context . The context is again derived from the clip-level context by temporal attention: .
The video-level OSTR models the long-term relationships between input object sequences in the whole video:
5 Answer Decoders
We follow the common settings for answer decoders (e.g., see ) which combine the final representation with the query using an MLP followed by a softmax to rank the possible answer choices. More details about the answer decoders per question types are available in the supplemental material. We use the cross-entropy as the loss function to training the model from end to end for all tasks except counting, where Mean Square Error is used.
Experiments
We evaluate our proposed HOSTR on the three public video QA benchmarks, namely, TGIF-QA , MSVD-QA and MSRVTT-QA . More details are as follows.
MSVD-QA consists of 50,505 QA pairs annotated from 1,970 short video clips. The dataset covers five question types: What, Who, How, When, and Where, of which 61% of the QA pairs for training, 13% for validation and 26% for testing.
MSRVTT-QA contains 10K real videos (65% for training, 5% for validation, and 30% for testing) with more than 243K question-answer pairs. Similar to MSVD-QA, questions are of five types: What, Who, How, When, and Where.
TGIF-QA is one of the largest Video QA datasets with 72K animated GIFs and 120K question-answer pairs. Questions cover four tasks - Action, Event Transition, FrameQA, Count. We refer readers to the supplemental material for the dataset description and statistics.
2 Comparison Against SOTAs
Implementation: We use Faster R-CNNhttps://github.com/airsplay/py-bottom-up-attention for frame-wise object detection. The number of object sequences per video for MSVD-QA , MSRVTT-QA is 40 and TGIF-QA is 50. We embed question words into 300-D vectors and initialize them with GloVe during training. Default settings are with 6 GCN layers for each OSTR unit. The feature dimension is set to be in all sub-networks.
We compare the performance of HOSTR against recent state-of-the-art (SOTA) methods on all three datasets. Prior results are taken from .
MSVD-QA and MSRVTT-QA: Table 1 shows detailed comparisons on MSVD-QA and MSRVTT-QA datasets. It is clear that our proposed method consistently outperforms all SOTA models. Specifically, we significantly improve performance on the MSVD-QA by 3.3 absolute points while the improvement on the MSRVTT-QA is more modest. As videos in the MSRVTT-QA are much longer (3 times longer than those in MSVD-QA) and contain more complicated interaction, it might require a larger number of input object sequences than what in our experiments (40 object sequences).
TGIF-QA: Table 2 presents the results on TGIF-QA dataset. As pointed out in , short-term motion features are helpful for action task while long-term motion features are crucial for event transition and count tasks. Hence, we provide two variants of our model: HOSTR (R) makes use of ResNet features as the context information for OSTR units; and HOSTR (R+F) makes use of the combination of ResNet features and motion features (extracted by ResNeXt, the same as in HCRN) as the context representation. HOSTR (R+F) shows exceptional performance on tasks related to motion. Note that we only use the context modulation at the video level to concentrate on the long-term motion. Even without the use of motion features, HOSTR (R) consistently shows more favorable performance than existing works.
The quantitative results prove the effectiveness of the object-oriented reasoning compared to the prior approaches of totally relying on frame-level features. Incorporating the motion features as context information also shows the flexibility of the design, suggesting that HOSTR can leverage a variety of input features and has the potentials to apply to other problems.
3 Ablation Studies
To provide more insight about our model, we examine the contributions of different design components to the model’s performance on the MSVD-QA dataset. We detail the results in Table 3. As shown, intra-object temporal attention seems to be more effective in summarizing the input object sequences to work with relational reasoning than BiLSTM. We hypothesize that it is due to the selective nature of the attention mechanism – it keeps only information relevant to the query.
As for the inter-object interaction, increasing the number of GCN layers up to 6 generally improves the performance. It gradually degrades when we stack more layers due to the gradient vanishing. The ablation study also points out the significance of the contextual representation as the performance steeply drops from 39.4 to 37.8 without them.
Last but not least, we conduct two experiments to demonstrate the significance of video hierarchical modeling. “1-level hierarchy” refers to when we replace all clip-level OSTRs with the global average pooling operation to summarize each object sequence into a vector while keeping the OSTR unit at the video level. “1.5-level hierarchy”, on the other hand, refers to when we use an average pooling operation at the video level while keeping the clip-level the same as in our HOSTR. Empirically, it shows that going deeper in hierarchy consistently improves performance on this dataset. The hierarchy may have greater effects in handling longer videos such as those in the MSRVTT-QA and TGIF-QA datasets.
4 Qualitative Analysis
To provide more analysis on the behavior of HOSTR in practice, we visualize the spatio-temporal graph formed during HOSTR operation on a sample in MSVD-QA dataset. In Fig. 4, the spatio-temporal graph of the two most important clips (judged by the temporal attention scores) are visualized in order of their appearance in the video’s timeline. Blue boxes indicate the six objects with highest edge weights (row summation of the adjacency matrix calculated in Eq.6). The red lines indicates the most prominent edges of the graph with intensity scaled to the edge strength.
In this example, HOSTR attended mostly on the objects related to the concepts relevant to answer the question (the girl, ball and dog). Furthermore, the relationships between the girl and her surrounding objects are the most important among the edges, and this intuitively agrees with how human might visually examine the scene given the question.
Conclusion
We presented a new object-oriented approach to Video QA where objects living in the video are treated as the primitive constructs. This brings us closer to symbolic reasoning, which is arguably more human-like. To realize this high-level idea, we introduced a general-purpose neural unit dubbed Object-oriented Spatio-Temporal Reasoning (OSTR). The unit reasons about its contextualized input – which is a set of object sequences – as instructed by the linguistic query. It first selectively transforms each sequence to an object node, then dynamically induces links between the nodes to build a graph. The graph enables iterative relational reasoning through collective refinement of object representation, gearing toward reaching an answer to the given query. The units are then stacked in a hierarchy that reflects the temporal structure of a typical video, allowing higher-order reasoning across space and time. Our architecture establishes new state-of-the-arts on major Video QA datasets designed for complex compositional questions, relational, temporal reasoning. Our analysis shows that object-oriented reasoning is a reliable, interpretable and effective approach to Video QA.
References
Appendix
Implementation Details
2 Answer decoders
In this subsection, we provide further details on how HOSTR predicts answers given the final representation and the query representation . Depending on question types, we design slightly different answer decoders. As for open-ended (MSVD-QA, MSRVTT-QA and FrameQA task in TGIF-QA) and multiple-choice questions (Action and Event Transition task in TGIF-QA), we utilize a classifier of 2-fully connected layers with the softmax function to rank words in a predefined vocabulary set or rank answer choices:
We use the cross-entropy as the loss function to training the model from end to end in this case. While prior works rely on hinge loss to train multiple-choice questions, we find that cross-entropy loss produces more favorable performance. For the count task in TGIF-QA dataset, we simply take the output of Eq. 12 and further feed it into a rounding operation for an integer output prediction. We use Mean Squared Error as the training loss for this task.
3 Training
In this paper, each video is split into clips of consecutive frames. This is done simply based on the empirical results. For the HOSTR (R+F) variant, we divided each video into eight clips of consecutive frames to fit in the predefined configuration of the ResNeXt pretrained modelhttps://github.com/kenshohara/video-classification-3d-cnn-pytorch. To mitigate the loss of temporal information when split a video up, there are overlapping parts between two video clips next to each other.
We train our model using Adam optimizer at an initial learning rate of and weight decay of . The learning rate is reduced after every 10 epochs for all tasks except the count task in TGIF-QA in which we reduce the learning rate by half after every 5 epochs. We use a batch size of 64 for MSVD-QA and MSRVTT-QA while that of TGIF-QA is 32. To be compatible with related works , we use accuracy as evaluation metric for multi-class classification tasks and MSE for the count task in TGIF-QA dataset..
All experiments are conducted on a single GPU NVIDIA Tesla V100-SXM2-32GB installed on a sever of 40 physical processors and 256 GB of memory running on Ubuntu 18.0.4. It may require a GPU of at least 16GB in order to accommodate our implementation. We use early stopping when validation accuracy decreases after 10 consecutive epochs. Otherwise, all experiments are terminated after 25 epochs at most. Depending on dataset sizes, it may take 9 hours (MSVD-QA) to 30 hours (MSRVTT-QA) of training. As for inference, it takes up to 40 minutes to complete (on MSRVTT-QA). HOSTR is implemented in Python 3.6 with Pytorch 1.2.0.
Dataset Description
is a relatively small dataset of 50,505 QA pairs annotated from 1,970 short video clips. The dataset covers five question types: What, Who, How, When, and Where, of which 61% of the QA pairs for training, 13% for validation and 26% for testing. We provide details on the number of question-answer pairs in Table 4.
contains 10K real videos (65% for training, 5% for validation, and 30% for testing) with more than 243K question-answer pairs. Similar to MSVD-QA, questions are of five types: What, Who, How, When, and Where. Details on the statistics of the MSRVTT-QA dataset is shown in Table 5.
is one of the largest Video QA datasets of 72K animated GIFs and 120K question-answer pairs. Questions cover four tasks - Action: multiple-choice task identifying what action repeatedly happens in a short period of time; Transition: multiple-choice task assessing if machines could recognize the transition between two events in a video; FrameQA: answers can be found in one of video frames without the need of temporal reasoning; and Count: requires machines to be able to count the number of times an action taking place. We take 10% the number of Q-A pairs and their associated videos in the training set per each task as the validation set. Details on the number of Q-A pairs and videos per split are provided in Table 6.
Qualitative Results
In addition to the example provided in the main paper, we present a few more examples of the spatio-temporal graph formed in the HOSTR model in Figure 5. In all examples presented, our HOSTR has successfully modeled the relationships between the targeted objects in the questions or their parts with surrounding objects. These examples clearly demonstrate the explainability and transparency of our model, explaining how the model arrives at the correct answers.
Complexity Analysis
We analyze the memory consumption on the number of object , hierarchy depth given that other factors like video length are fixed; and OSTR units in one layer share parameters.
There are two types of memory is from hidden unit activations, is for parameters and related terms. We found that is linear with , constant to . On the other hand, is linear with and constant to . These figures are on-par with commonly used sequential models. No major extra memory consumption is introduced.