Hierarchical Conditional Relation Networks for Video Question Answering

Thao Minh Le, Vuong Le, Svetha Venkatesh, Truyen Tran

Introduction

Answering natural questions about a video is a powerful demonstration of cognitive capability. The task involves acquisition and manipulation of spatio-temporal visual representations guided by the compositional semantics of the linguistic cues . As questions are potentially unconstrained, VideoQA requires deep modeling capacity to encode and represent crucial video properties such as object permanence, motion profiles, prolonged actions, and varying-length temporal relations in a hierarchical manner. For VideoQA, the visual representations should ideally be question-specific and answer-ready.

The current approach toward modeling videos for QA is to build neural architectures in which each sub-system is either designed for a specific tailor-made purpose or for a particular data modality. Because of this specificity, such hand crafted architectures tend to be non-optimal for changes in data modality , varying video length or question types (such as frame QA versus action count ). This has resulted in proliferation of heterogeneous networks.

In this work we propose a general-purpose reusable neural unit called Conditional Relation Network (CRN)\text{CRN}) that encapsulates and transforms an array of objects into a new array conditioned on a contextual feature. The unit computes sparse high-order relations between the input objects, and then modulates the encoding through a specified context (See Fig. 2). The flexibility of CRN and its encapsulating design allow it to be replicated and layered to form deep hierarchical conditional relation networks (HCRN) in a straightforward manner. The stacked units thus provide contextualized refinement of relational knowledge from video objects – in a stage-wise manner it combines appearance features with clip activity flow and linguistic context, and follows it by integrating in context from the whole video motion and linguistic features. The resulting HCRN is homogeneous, agreeing with the design philosophy of networks such as InceptionNet , ResNet and FiLM .

The hierarchy of the CRNs are as follows – at the lowest level, the CRNs encode the relations between frame appearance in a clip and integrate the clip motion as context; this output is processed at the next stage by CRNs that now integrate in the linguistic context; in the following stage, the CRNs capture the relation between the clip encodings, and integrate in video motion as context; in the final stage the CRN integrates the video encoding with the linguistic feature as context (See Fig. 3). By allowing the CRNs to be stacked hierarchically, the model naturally supports modeling hierarchical structures in video and relational reasoning; by allowing appropriate context to be introduced in stages, the model handles multimodal fusion and multi-step reasoning. For long videos further levels of hierarchy can be added enabling encoding of relations between distant frames.

We demonstrate the capability of HCRN in answering questions in major VideoQA datasets. The hierarchical architecture with four-layers of CRN units achieves favorable answer accuracy across all VideoQA tasks. Notably, it performs consistently well on questions involving either appearance, motion, state transition, temporal relations, or action repetition demonstrating that the model can analyze and combine information in all of these channels. Furthermore HCRN scales well on longer length videos simply with the addition of an extra layer. Fig. 1 demonstrates several representative cases those were difficult for the baseline of flat visual-question interaction but can be handled by our model.

Our model and results demonstrate the impact of building general-purpose neural reasoning units that support native multimodality interaction in improving robustness and generalization capacities of VideoQA models.

Related Work

Our proposed HCRN model advances the development of VideoQA by addressing two key challenges: (1) Efficiently representing videos as amalgam of complementing factors including appearance, motion and relations, and (2) Effectively allows the interaction of such visual features with the linguistic query.

Spatio-temporal video representation is traditionally done by variations of recurrent networks (RNNs) among which many were used for VideoQA such as recurrent encoder-decoder , bidirectional LSTM and two-staged LSTM . To increase the memorizing ability, external memory can be added to these networks . This technique is more useful for videos that are longer and with more complex structures such as movies and TV programs with extra accompanying channels such as speech or subtitles. On these cases, memory networks were used to store multimodal features for later retrieval. Memory augmented RNNs can also compress video into heterogenous sets of dual appearance/motion features. While in RNNs, appearance and motion are modeled separately, 3D and 2D/3D hybrid convolutional operators intrinsically integrates spatio-temporal visual information and are also used for VideoQA . Multiscale temporal structure can be modeled by either mixing short and long term convolutional filters or combining pre-extracted frame features non-local operators . Within the second approach, the TRN network demonstrates the role of temporal frame relations as an another important visual feature for video reasoning and VideoQA . Relations of predetected objects were also considered in a separate processing stream and combined with other modalities in late-fusion . Our HCRN model emerges on top of these trends by allowing all three channels of video information namely appearance, motion and relations to iteratively interact and complement each other in every step of a hierarchical multi-scale framework.

Earlier attempts for generic multimodal fusion for visual reasoning includes bilinear operators, either applied directly or through attention . While these approaches treat the input tensors equally in a costly joint multiplicative operation, HCRN separates conditioning factors from refined information, hence it is more efficient and also more flexible on adapting operators to conditioning types.

Temporal hierarchy has been explored for video analysis , most recently with recurrent networks and graph networks . However, we believe we are the first to consider hierarchical interaction of multi-modalities including linguistic cues for VideoQA.

Linguistic query–visual feature interaction in VideoQA has traditionally been formed as a visual information retrieval task in a common representation space of independently transformed question and referred video . The retrieval is more convenient with heterogeneous memory slots . On top of information retrieval, co-attention between the two modalities provides a more interactive combination . Developments along this direction include attribute-based attention , hierarchical attention , multi-head attention , multi-step progressive attention memory or combining self-attention with co-attention . For higher order reasoning, question can interact iteratively with video features via episodic memory or through switching mechanism . Multi-step reasoning for VideoQA is also approached by and with refined attention.

Unlike these techniques, our HCRN model supports conditioning video features with linguistic clues as a context factor in every stage of the multi-level refinement process. This allows linguistic cue to involve earlier and deeper into video presentation construction than any available methods.

Neural building blocks - Beyond the VideoQA domain, CRN unit shares the idealism of uniformity in neural architecture with other general purpose neural building blocks such as the block in InceptionNet , Residual Block in ResNet , Recurrent Block in RNN, conditional linear layer in FiLM , and matrix-matrix-block in neural matrix net . Our CRN departs significantly from these designs by assuming an array-to-array block that supports conditional relational reasoning and can be reused to build networks of other purposes in vision and language processing.

Method

where θ\theta is the model parameters of scoring function F\mathcal{F}.

With these representations, we now describe our new hierarchical architecture for VideoQA (see Fig. 3). We first present the core compositional computation unit that serves as building blocks for the architecture in Section 3.1. In the following sub-section, we propose to design F\mathcal{F} as a layer-by-layer network architecture that can be built by simply stacking the core units in a particular manner.

1 Conditional Relation Network Unit

When in use for VideoQA, CRN’s input array is composed of features at either frame or short-clip levels. The objects {si}i=1n\{s_{i}\}_{i=1}^{n} greatly share mutual information and it is redundant to consider all possible combinations of given objects. Therefore, applying a sampling scheme on the set of subsets (line 4 of Alg. 1) is crucial for redundancy reduction and computational efficiency. We borrow the sampling trick in to build sets of tt selected subsets QselectedkQ_{\textrm{selected}}^{k}. Regarding the choice of kmaxk_{\text{max}}, we choose kmax=n−1k_{\text{max}}=n-1 in later experiments, resulting in the output array of size n−2n-2 if n>2n>2 and array of size 11 if n=2n=2.

As a choice in implementation, the functions gk(.),pk(.)g^{k}(.),p^{k}(.) are simple average-pooling. In generic form, they can be any aggregation sub-networks that join a random set into a single representation. Meanwhile, hk(.,.)h^{k}(.,.) is a MLP running on top of feature concatenation that models the non-linear relationships between multiple input modalities. We tie parameters of the conditioning sub-network hk(.,.)h^{k}(.,.) across the subsets of the same size kk. In our implementation, hk(.,.)h^{k}(.,.) consists of a single linear transformation followed by an ELU activation.

It may be of concern that the relation formed by a particular subset may be unnecessary to model kk-tuple relations, we optionally design a self-gating mechanism similar to to regulate the feature flow to go through each CRN module. Formally, the conditioning function hk(.,.)h^{k}(.,.) in that case is given by:

where [.,.][.,.] denotes the tensor concatenation, σ\sigma is sigmoid function, and Wh1,Wh2W_{h_{1}},W_{h_{2}} are linear weights.

2 Hierarchical Conditional Relation Networks

We use CRN blocks to build a deep network architecture to exploit inherent characteristics of a video sequence namely temporal relations, motion, and the hierarchy of video structure, and to support reasoning guided by linguistic questions. We term the proposed network architecture Hierarchical Conditional Relation Networks (HCRN) (see Fig. 3 ). The design of the HCRN by stacking reusable core units is partly inspired by modern CNN network architectures, of which InceptionNet and ResNet are the most well-known examples.

A model for VideoQA should distill the visual content in the context of the question, given the fact that much of the visual information is usually not relevant to the question. Drawing inspiration from the hierarchy of video structure, we boil down the problem of VideoQA into a process of video representation in which a given video is encoded progressively at different granularities, including short clip (consecutive frames) and entire video levels. It is crucial that the whole process conditions on linguistic cue. In particular, at each hierarchy level, we use two stacked CRN units, one conditioned on motion features followed by one conditioned on linguistic cues. Intuitively, the motion feature serves as a dynamic context shaping the temporal relations found among frames (at the clip level) or clips (at the video level). As the shaping effect is applied to all relations, self-gating is not needed, and thus a simple MLP suffices. On the other hand, the linguistic cues are by nature selective, that is, not all relations are equally relevant to the question. Thus we utilize the self-gating mechanism in Eq. (2) for the CRN units which condition on question representation.

With this particular design of network architecture, the input array at clip level consists of frame-wise appearance feature vectors {v^ij}\{\hat{v}_{ij}\}, while that at a video level is the output at the clip level. Meanwhile, the motion conditioning feature at clip level CRNs are corresponding clip motion feature vector f^i\hat{f}_{i}. They are further passed to an LSTM, whose final state is used as video-level motion features. Note that this particular implementation is not the only option. We believe we are the first to progressively incorporate multiple modalities of input in such a hierarchical manner in contrast to the typical approach of treating appearance features and motion features as a two-stream network.

To handle a long video of thousand frames, which is equivalent to dozens of short-term clips, there are two options to reduce the computational cost of CRN in handling large sets of subsets {Qk∣k=2,3,...,kmax}\{Q^{k}|k=2,3,...,k_{\text{max}}\} given an input array SS: limit the maximum subset size kmaxk_{\text{max}} or extend the HCRN to deeper hierarchy. For the former option, this choice of sparse sampling may have potential to lose critical relation information of specific subsets. The latter, on the other hand, is able to densely sample subsets for relation modeling. Specifically, we can group NN short-term clips into N1×N2N_{1}\times N_{2} hyper-clips, of which N1N_{1} is the number of the hyper-clips and N2N_{2} is the number of short-term clips in one hyper-clip. By doing this, our HCRN now becomes a 3-level of hierarchical network architecture.

where, [.,.]\left[.,.\right] denotes concatenation operation, and ⊙\odot is the Hadamard product.

3 Answer Decoders and Loss Functions

The cross-entropy is used as the loss function.

For repetition count task, we use a linear regression function taking y′y^{\prime} in Eq. (8) as input, followed by a rounding function for integer count results. The loss for this task is Mean Squared Error (MSE).

We use the popular hinge loss of pairwise comparisons, max(0,1+sn−sp)\text{max}\left(0,1+s^{n}-s^{p}\right), between scores for incorrect sns^{n} and correct answers sps^{p} to train the network.

4 Complexity Analysis

We provide a brief analysis here, leaving detailed derivations in Supplement. For a fixed sampling resolution tt, a single forward pass of CRN would take quadratic time in kmaxk_{\text{max}}. For an input array of length nn, feature size FF, the unit produces an output array of size kmax−1k_{\text{max}}-1 of the same feature dimensions. The overall complexity of HCRN depends on design choice for each CRN unit and specific arrangement of CRN units. For clarity, let t=2t=2 and kmax⁡=n−1k_{\max}=n-1, which are found to work well in later experiments. Suppose there are NN clips of length TT, making a video of length L=NTL=NT. A 2-level architecture of Fig. 3 needs 2TLF2TLF time to compute the CRNs at the lowest level, and 2NLF2NLF time to compute the second level, totaling 2(T+N)LF2(T+N)LF time.

Let us now analyze a 3-level architecture that generalizes the one in Fig. 3. The NN clips are organized into MM sub-videos, each has QQ clips, i.e., N=MQN=MQ. The clip-level CRNs remain the same. At the next level, each sub-video CRN takes as input an array of length QQ, whose elements have size (T−4)F(T-4)F. Using the same logic as before, the set of sub-video-level CRNs cost 2NMLF2\frac{N}{M}LF time. A stack of two sub-video CRNs now produces an output array of size (Q−4)(T−4)F(Q-4)(T-4)F, serving as an input object in an array of length MM for the video-level CRNs. Thus the video-level CRNs cost 2MLF2MLF. Thus the total cost for 3-level HCRN is in the order of 2(T+NM+M)LF2(T+\frac{N}{M}+M)LF.

Compared to the 2-level HCRN, the a 3-level HCRN reduces computation time by 2(N−NM−M)LF≈2NLF2(N-\frac{N}{M}-M)LF\approx 2NLF assuming N≫max⁡{M,NM}N\gg\max\left\{M,\frac{N}{M}\right\}. As N=LTN=\frac{L}{T}, this reduces to 2NLF=2L2TF2NLF=2\frac{L^{2}}{T}F. In practice TT is often fixed, thus the saving scales quadratically with video length LL, suggesting that hierarchy is computational efficient for long videos.

Experiments

This is currently the most prominent dataset for VideoQA, containing 165K QA pairs and 72K animated GIFs. The dataset covers four tasks addressing unique properties of video. Of which, the first three require strong spatio-temporal reasoning abilities: Repetition Count - to retrieve number of occurrences of an action, Repeating Action- multi-choice task to identify the action repeated for a given number of times, State Transition - multi-choice tasks regarding temporal order of events. The last task - Frame QA - is akin to image QA where a particular frame in a video is sufficient to answer the questions.

This is a small dataset of 50,505 question answer pairs annotated from 1,970 short clips. Questions are of five types, including what, who, how, when and where.

The dataset contains 10K videos and 243K question answer pairs. Similar to MSVD-QA, questions are of five types. Compared to the other two datasets, videos in MSRVTT-QA contain more complex scenes. They are also much longer, ranging from 10 to 30 seconds long, equivalent to 300 to 900 frames per video.

We use accuracy to be the evaluation metric for all experiments, except those for repetition count on TGIF-QA dataset where Mean Square Error (MSE) is applied.

2 Implementation Details

Videos are segmented into 8 clips, each clip contains 16 frames by default. Long videos in MSRVTT-QA are additionally segmented into 24 clips for evaluating the ability of handling very long sequences. Unless otherwise stated, the default setting is with a 2-level HCRN as depicted in Fig. 3, and d=512d=512, t=1t=1. We train the model initially at learning rate of 10−410^{-4} and decay by half after every 10 epochs. All experiments are terminated after 25 epochs and reported results are at the epoch giving the best validation accuracy. Pytorch implementation of the model is available onlinehttps://github.com/thaolmk54/hcrn-videoqa.

3 Results

We compare our proposed model with state-of-the-art methods (SoTAs) on aforementioned datasets. For TGIF-QA, we compare with most recent SoTAs, including , over four tasks. These works, except for , make use of motion features extracted from optical flow or 3D CNNs.

The results are summarized in Table 2 for TGIF-QA, and in Fig. 4 for MSVD-QA and MSRVTT-QA. Reported numbers of the competitors are taken from the original papers and . It is clear that our model consistently outperforms or is competitive with SoTA models on all tasks across all datasets. The improvements are particularly noticeable when strong temporal reasoning is required, i.e., for the questions involving actions and transitions in TGIF-QA. These results confirm the significance of considering both near-term and far-term temporal relations toward finding correct answers.

The MSVD-QA and MSRVTT-QA datasets represent highly challenging benchmarks for machine compared to the TGIF-QA, thanks to their open-ended nature. Our model HCRN outperforms existing methods on both datasets, achieving 36.1% and 35.6% accuracy which are 1.7 points and 0.6 points improvement on MSVD-QA and MSRVTT-QA, respectively. This suggests that the model can handle both small and large datasets better than existing methods.

Finally, we provide a justification for the competitive performance of our HCRN against existing rivals by comparing model features in Table 3. Whilst it is not straightforward to compare head-to-head on internal model designs, it is evident that effective video modeling necessitates handling of motion, temporal relation and hierarchy at the same time. We will back this hypothesis by further detailed studies in Section 4.3.2 (for motion, temporal relations, shallow hierarchy) and Section 4.3.3 (deep hierarchy).

3.2 Ablation Studies

To provide more insight about our model, we conduct extensive ablation studies on TGIF-QA with a wide range of configurations. The results are reported in Table 4. Full 22-level HCRN denotes the full model of Fig. 3 with kmax=n−1,t=2k_{max}=n-1,t=2. Overall we find that ablating any of design components or CRN units would degrade the performance for temporal reasoning tasks (actions, transition and action counting). The effects are detailed as follows.

Without relations (kmax=1k_{max}=1) the performance drops significantly on actions and events reasoning. This is expected since those questions often require putting actions and events in relation with a larger context (e.g., what happens before something else). In this case, the frame QA benefits more from increasing sampling resolution tt because of better chance to find a relevant frame. However, when taking relations into account (kmax>1k_{max}>1), we find that HCRN is robust against sampling resolution tt but depends critically on the maximium relation order kmaxk_{max}. The relative independence w.r.t. tt can be due to visual redundancy between frames, so that resampling may capture almost the same information. On the other hand, when considering only low-order object relations, the performance is significantly dropped in all tasks, except frame QA. These results confirm that high-order relations are required for temporal reasoning. As the frame QA task requires only reasoning on a single frame, incorporating temporal information might confuse the model.

We design two simpler models with only one CRN layer: ▶\blacktriangleright 11-level, 11 CRN video on key frames only: Using only one CRN at the video-level whose input array consists of key frames of the clips. Note that video-level motion features are still maintained. ▶\blacktriangleright 1.51.5-level, clip CRNs →\rightarrow pooling: Only the clip-level CRNs are used, and their outputs are mean-pooled to represent video. The pooling operation represents a simplistic relational operation across clips. The results confirm that a hierarchy is needed for high performance on temporal reasoning tasks.

We evaluate the following settings: ▶\blacktriangleright w/o short-term motions: Remove all CRN units that condition on the short-term motion features (clip level) in the HCRN. ▶\blacktriangleright w/o long-term motions: Remove the CRN unit that conditions on the long-term motion features (video level) in the HCRN. ▶\blacktriangleright w/o motions: Remove motion feature from being used by HCRN. We find that motion, in agreeing with prior arts, is critical to detect actions, hence computing action count. Long-term motion is particularly significant for counting task, as this task requires maintaining global temporal context during the entire process. For other tasks, short-term motion is usually sufficient. E.g. in action task, wherein one action is repeatedly performed during the entire video, long-term context contributes little. Not surprisingly, motion does not play the positive role in answering questions on single frames as only appearance information needed.

Linguistic cues represent a crucial context for selecting relevant visual artifacts. For that we test the following ablations: ▶\blacktriangleright w/o quest.@clip level: Remove the CRN unit that conditions on question representation at clip level. ▶\blacktriangleright w/o quest.@video level: Remove the CRN unit that conditions on question representation at video level. ▶\blacktriangleright w/o linguistic condition: Exclude all CRN units conditioning on linguistic cue while the linguistic cue is still in the answer decoder. Likewise, gating offers a selection mechanism. Thus we study its effect as follows: ▶\blacktriangleright wo/ gate: Turn off the self-gating mechanism in all CRN units. ▶\blacktriangleright w/ gate quest. & motion: Turn on the self-gating mechanism in all CRN units.

We find that the conditioning question provides an important context for encoding video. Conditioning features (motion and language), through the gating mechanism in Eq. (2), offers further performance gain in action and counting tasks, possibly by selectively passing question-relevant information up the inference chain.

3.3 Deepening model hierarchy

We test the scalability of the HCRN on long videos in the MSRVTT-QA dataset, which are organized into 24 clips (3 times longer than other two datasets). We consider two settings: ▶\blacktriangleright 22-level hierarchy, 2424 clips→\rightarrow11 vid: The model is as illustrated in Fig. 3, where 24 clip-level CRNs are followed by a video-level CRN. ▶\blacktriangleright 33-level hierarchy, 2424 clips→\rightarrow44 sub-vids→\rightarrow11 vid: Starting from the 24 clips as in the 22-level hierarchy, we group 24 clips into 4 sub-videos, each is a group of 6 consecutive clips, resulting in a 33-level hierarchy. These two models are designed to have similar number of parameters, approx. 50M.

The results are reported in Table 5. Unlike existing methods which usually struggle with handling long videos, our method is scalable for them by offering deeper hierarchy, as analyzed theoretically in Section 3.4. Using a deeper hierarchy is expected to significantly reduce the training time and inference time for HCRN, especially when the video is long. In our experiments, we achieve 4 times reduction in training and inference time by going from 2-level HCRN to 3-level counterpart whilst maintaining the same performance.

Discussion

We introduced a general-purpose neural unit called Conditional Relational Networks (CRNs) and a method to construct hierarchical networks for VideoQA using CRNs as building blocks. A CRN is a relational transformer that encapsulates and maps an array of tensorial objects into a new array of the same kind, conditioned on a contextual feature. In the process, high-order relations among input objects are encoded and modulated by the conditioning feature. This design allows flexible construction of sophisticated structure such as stack and hierarchy, and supports iterative reasoning, making it suitable for QA over multimodal and structured domains like video. The HCRN was evaluated on multiple VideoQA datasets (TGIF-QA, MSVD-QA, MSRVTT-QA) demonstrating competitive reasoning capability.

Different to temporal attention based approaches which put effort into selecting objects, HCRN concentrates on modeling relations and hierarchy in video. This difference in methodology and design choices leads to distinctive benefits. CRN units can be further augmented with attention mechanisms to cover better object selection ability, so that related tasks such as frame QA can be further improved.

The examination of CRN in VideoQA highlights the importance of building generic neural reasoning unit that supports native multimodal interaction in improving robustness of visual reasoning. We wish to emphasize that the unit is general-purpose, and hence is applicable for other reasoning tasks, which we will explore. These includes an extension to consider the accompanying linguistic channels which are crucial for TVQA and MovieQA tasks.

References